Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries
cs.CL, cs.AI
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 13 pages, 1 figure, 5 tables. Code, full per-request logs, and all judge votes: https://github.com/adorosario/why-rags-hallucinate
Code: https://github.com/adorosario/why-rags-hallucinate
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Ragas: Automated Evaluation of Retrieval Augmented Generation
- SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
- Language Models (Mostly) Know What They Know
- Why Language Models Hallucinate
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
- LLM Evaluators Recognize and Favor Their Own Generations
- ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
- Measuring short-form factuality in large language models
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering