LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes and What Recovers It
cs.CL, cs.AI
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/composo-ai/omission-bench
Terminology
Sources
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- CARE: A Conformal Safety Layer for Medical Summarization
- Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets
- Same Verdict, Different Reasons: LLM-as-a-Judge and Clinician Disagreement on Medical Chatbot Completeness
- AbsenceBench: Language Models Can't Tell What's Missing
- Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses
- Extrinsically-Focused Evaluation of Omissions in Medical Summarization
- Beyond Literal Summarization: Redefining Hallucination for Medical SOAP Note Evaluation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering