Building Legal Reward Models for Grounding and Abstention
| Source: arXiv AI
Tags: reward models, legal AI, RAG, DPO, LegalRewardBench, grounding
LegalRewardBench exposes that reward models optimized for general preferences fail in legal RAG settings: contextual DPO with length-balanced training data improves grounded legal evaluation by up to 25.6pp, and cross-jurisdiction transfer from Victorian criminal law boosts US housing statute QA by 16.2pp.
Details
Legal AI systems operating with retrieval-augmented generation must do two hard things: ground reasoning in retrieved evidence and refuse to answer when evidence is insufficient. Standard reward models, trained on general preference data, are poorly equipped to evaluate these behaviors.\n\nPublished at the ICML 2026 AI4Law Workshop, this paper introduces a framework for converting existing legal QA datasets into contextual preference data and uses it to build LegalRewardBench (LRB), covering in-distribution inputs, hard-benign prompts, obfuscated queries, and domain/structural distribution shifts.\n\nFindings show contextual DPO improves grounded evaluation, but results are sensitive to how preference data is constructed. Length-balanced augmentation combining legal and general contextual data yields the strongest gains — up to +25.6pp over baseline on grounded legal evaluation. Notably, models refined on Victorian criminal-law data transfer to US legal benchmarks, improving Housing Statute QA by +16.2pp.\n\nFor enterprise legal AI teams, this provides concrete evidence that reward model fine-tuning with domain-appropriate, length-balanced preference data is necessary for reliable RAG grounding and abstention — and that cross-jurisdiction transfer reduces the data collection burden.