BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law
cs.CL, cs.AI
Submitted: 2026-05-27
Updated: 2026-08-31
Code: https://github.com/SebastianNagl/benger-platform
Terminology
Sources
- Legal RAG Bench: an end-to-end benchmark for legal RAG
- The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
- GreekBarBench: A Challenging Benchmark for Free-Text Legal Reasoning and Citations
- Better Call CLAUSE: A Discrepancy Benchmark for Auditing LLMs Legal Reasoning Capabilities
- LEXam: Benchmarking Legal Reasoning on 340 Law Exams
- AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark
- LLM Evaluators Recognize and Favor Their Own Generations
- Ready Jurist One: Benchmarking Language Agents for Legal Intelligence in Dynamic Environments
- LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain
- PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
- BERTScore: Evaluating Text Generation with BERT
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering