From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers
| Source: Apple ML Research
Tags: Apple, reward modeling, RLHF, question answering, RAG, post-training, alignment
Apple researchers introduce a rubric-based reward framework for open-domain QA that improves over instruction-tuned baselines by 6.5%, using query-specific rubrics grounded in retrieved evidence across three quality dimensions — composition, grounding, and instruction-following.
Details
Apple ML Research has published a paper introducing rubric-based reward modeling for grounded knowledge question answering, targeting a core weakness in standard RLHF: holistic scalar rewards fail to capture the multiple orthogonal quality dimensions that good answers require simultaneously. The framework generates query-specific rubrics conditioned on retrieved evidence rather than generic quality heuristics. These rubrics are decomposed into multiple quality dimensions — composition, grounding, and instruction-following — providing fine-grained supervision signals during post-training rather than a single averaged score. Results show a 6.5% average improvement over instruction-tuned baselines and a 4% gain over flat (non-decomposed) rubric variants, with consistent performance across all evaluated datasets. The paper isolates two key contributions: conditioning on retrieved documents improves factual grounding, while decomposing rubrics into dimension-specific sub-scores further improves coherence and organization. For teams building RAG pipelines with RLHF fine-tuning, this approach offers a practical path to richer reward signals without requiring human preference labels at the granularity of individual quality dimensions — the rubrics are generated automatically per query.