Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering
cs.CL
Submitted: 2026-09-14
Updated: 2026-09-17
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) have been widely adopted for clinical question answering (QA).
Terminology
Abstract
Large language models (LLMs) have been widely adopted for clinical question answering (QA). Current systems can attach citations to their answers, but these often point to broad texts, leaving time-pressed clinicians unable to verify them efficiently. An alternative is to ensure that responses are verifiable by construction: providing fine-grained verbatim quotes from reference material that substantiate claims, so users can verify an answer without opening other documents. In this paper, we evaluate the ability of current models to perform this task end-to-end: from providing citations for every factual claim, to producing verbatim quotes, to ensuring that those quotes fully substantiate the claims. To do so, we build a standardized harness over four clinical practice guidelines and evaluate twelve LLMs on 222 synthetic clinical questions, measuring each of these stages separately. We find that most models can attach verbatim quotes to over 90% of their claims from prompting alone, apart from some lightweight models such as claude-haiku-4.5. Yet these quotes often fail to substantiate every detail of the claims they accompany. For instance, claude-opus-5 produces verbatim quotes for 98.0% of its claims, but fully substantiates only 37.1%. Our work provides insights into the current capability gap of LLMs in building verifiable clinical QA systems, along with artifacts for future research.
Sources
- Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models
- Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Teaching language models to support answers with verified quotes
- WebGPT: Browser-assisted question-answering with human feedback
- DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
- Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering