Building Legal Reward Models for Grounding and Abstention
cs.CL, cs.AI
Submitted: 2026-09-13
Updated: 2026-09-13
Comments: Published at ICML 2026 AI4Law Workshop
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models are increasingly used in high-stakes domains such as law, where systems must ground their reasoning in retrieved evidence and abstain when that evidence is insufficient.
Terminology
Abstract
Large language models are increasingly used in high-stakes domains such as law, where systems must ground their reasoning in retrieved evidence and abstain when that evidence is insufficient. However, existing reward models are largely optimised for general preferences rather than contextual grounding, limiting their ability to evaluate these behaviours in retrieval-augmented generation (RAG) settings. We introduce a framework for transforming existing legal QA datasets into contextual preference data and use it to construct LegalRewardBench (LRB), a benchmark for evaluating grounded legal generation under noisy and insufficient retrieval conditions. Across general and legal contextual evaluation, we find that contextual DPO improves grounded evaluation, but performance is sensitive to preference-data construction. Length-balanced augmentation substantially improves grounded legal evaluation, with the strongest configuration combining length-balanced legal and general contextual preference data and improving performance by up to+25.6 pp over baseline. We further find evidence of cross-jurisdiction transfer: models contextually refined primarily on Victorian criminal-law data improve grounded evaluation on external US legal benchmarks, including a+16.2 pp improvement on Housing Statute QA. Together, these results provide a reproducible foundation for constructing and evaluating grounded legal reward models in retrieval-augmented settings.
Sources
- Atla Selene Mini: A General Purpose Evaluation Model
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Constitutional AI: Harmlessness from AI Feedback
- Legal RAG Bench: an end-to-end benchmark for legal RAG
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- RLHF Workflow: From Reward Modeling to Online RLHF
- Natural Language Processing in the Legal Domain
- GLM-5: from Vibe Coding to Agentic Engineering
- The Llama 3 Herd of Models
- Ministral 3
- Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
- GPT-4 Technical Report
- Gemma 3 Technical Report
- Gemma 2: Improving Open Language Models at a Practical Size
- LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain
- Qwen2.5 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering