Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization
cs.CL
Submitted: 2026-01-07
Updated: 2026-08-26
Comments: Accepted at EMNLP 2026 Main; webpage at https://yao-dou.github.io/gavel/
Project page: https://yao-dou.github.io/gavel
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) now support contexts of up to 1M tokens, but their strengths and weaknesses on complex long-context tasks remain unclear.
Terminology
Abstract
Large language models (LLMs) now support contexts of up to 1M tokens, but their strengths and weaknesses on complex long-context tasks remain unclear. To study this, we focus on multi-document legal case summarization, where a single case often spans many documents exceeding 100K tokens. We systematically evaluate 12 frontier LLMs with Gavel, which consists of Gavel-Ref, a reference-based evaluation framework with checklist, residual-fact, and writing-style evaluations, and Gavel-Agent, a reference-free agent for evaluating factual coverage directly from source documents. Our results show that current models are more prone to omitting key information than hallucinating. They all perform well on simple checklist items, such as filing date, but struggle with rare and complex items, such as settlements. Performance also declines as case length increases. To meta-evaluate Gavel, we collect 160 hours of human annotations. Gavel-Agent reduces token usage by at least 36% compared to end-to-end and chunk-by-chunk methods while achieving competitive performance. Gavel-Agent also generalizes to the medical domain, performing the best with at least 77% fewer tokens.
Sources
- GPT-4 Technical Report
- gpt-oss-120b & gpt-oss-20b Model Card
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Walking Down the Memory Maze: Beyond Context Limit through Interactive Reading
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- GraphReader: Building Graph-based Agent to Enhance Long-Context Abilities of Large Language Models
- Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
- A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
- Text Summarization with Pretrained Encoders
- CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions
- OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning
- OpenAgents: An Open Platform for Language Agents in the Wild
- Qwen3 Technical Report
- Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution
- HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly
- ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists
- Legal Case Document Summarization: Extractive and Abstractive Methods and their Evaluation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering