CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence".
Jane: The gist The CiteVQA benchmark introduces an evaluation framework that requires models to return element-level bounding-box citations alongside each answer, evaluating both jointly.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, the CiteVQA paper sets up this whole evaluation framework where models have to return element-level bounding-box citations alongside their answers. They are using a metric called Strict Attributed Accuracy, or SAA, to judge both the answer and the cited region at the same time.
Jane: The core of their claim is that this joint evaluation exposes what they call Attribution Hallucination, which is when a model gives you the right answer but grounds it in completely incorrect visual evidence.
Lu: They built this benchmark around one thousand eight hundred ninety-seven questions spread across seven hundred eleven PDFs covering seven different domains and two languages, with documents averaging about forty point six pages each to keep it realistic <ref:2605.12882#pg1>.
Meng: The paper claims that by forcing models to provide these precise citations, they can establish a rigorous standard for measuring evidence fidelity in high-stakes areas like finance or law.
Lalam: It really matters because this isn't just about getting the answer right; it’s about knowing the specific visual source behind that claim, which is what makes an AI trustworthy for critical tasks.
Conclusion: Tom: So, looking at the CiteVQA benchmark, we see it’s a direct response to a weakness in how we evaluate document understanding before this work came out by Dongsheng Ma and his team.
Jane: The authors are essentially arguing that answer-only scoring is insufficient because it hides this specific failure mode where the model mixes up correct text with wrong visual evidence.
Lu: The implication here for the wider field is that we need to move beyond just scoring the final output and start demanding verifiable, element-level grounding from multimodal models.
Meng: For practical application, this means any system we build that reads documents needs this kind of accountability to ensure it’s not accidentally citing something irrelevant just because it sounds plausible.
Lalam: It settles a big question: can we build reliable document intelligence if we require every single claim to point directly to a specific location on the page?
Tom: Exactly, and the paper shows that while some models do well on answering things, their SAA scores are much lower than what's required for high-stakes work.
Jane: So, CiteVQA gives us the tools to test these systems under real conditions and show where they actually fall short in terms of accurate attribution.
Peking University
cs.CL, cs.CV
Submitted: 2026-05-13
Updated: 2026-10-08
Code: https://github.com/opendatalab/CiteVQA
Importance score: 90/100
The gist: The gist The CiteVQA benchmark introduces an evaluation framework that requires models to return element-level bounding-box citations alongside each answer, evaluating both jointly.
Key concepts
- CiteVQA Benchmark
- A new evaluation framework that requires AI models to return specific, element-level bounding-box citations alongside their answers. It tests whether the model can link its generated text directly to the exact visual source within a document.
- Strict Attributed Accuracy (SAA)
- The core metric used to judge model performance. A prediction only scores as correct if both the answer provided and the cited region of evidence are accurate. This ensures that models aren't just guessing; they must correctly identify and link verifiable facts.
- Attribution Hallucination
- A critical vulnerability discovered in models where they generate a correct answer but point to entirely incorrect visual evidence within the document. This phenomenon exposes a gap between linguistic understanding and accurate spatial grounding, showing models can be wrong about their sources even when their text is right.
Terminology
Summary
The gist The CiteVQA benchmark introduces an evaluation framework that requires models to return element-level bounding-box citations alongside each answer, evaluating both jointly.
How it works
-
CiteVQA comprises 1,897 questions across 711 PDFs spanning seven domains and two languages, averaging 40.6 pages per document.
-
The core evaluation is Strict Attributed Accuracy (SAA), which credits a prediction only when the answer and the cited region are both correct.
-
The benchmark mandates that models provide the precise PDF source supporting their answer at the granularity of element-level bounding-box citations, ensuring every generated claim is visually verifiable by human users.
-
A scalable, automated annotation pipeline was developed to generate ground-truth citations by identifying crucial evidence via masking ablation and validating them through expert review.
Key Findings on Model Performance
**- Auditing 20 MLLMs reveals a pervasive Attribution Hallucination, where models frequently produce the right answer while citing the wrong region. The strongest system, Gemini-3.1-Pro-Preview, achieves an SAA of only 76.0, and the strongest open-source MLLM reaches just 22.5. This exposes a reliability gap that answer-only evaluations overlook. The performance disparity shows that while models possess the perceptual capacity to extract information for a correct answer, they lack the ability to precisely link that information to its specific spatial source within the document. 5.3 Case Study To intuitively illustrate the disparity between linguistic performance and attribution accuracy—specifically why some powerful MLLMs achieve high Ans. but low SAA scores—we provide a case study in Figure 7. While Qwen3-VL-235B-A22B answers correctly (Ans.=1), it yields SAA=0 because its evidence crops are either blank or incomplete. 5.4 Impact of Multi-document Complexity The challenge of attribution is significantly exacerbated as the environment shifts from Single-Doc to Multi-Doc (N-Gold) settings. For instance, Gemini-3.1-Pro’s Recall drops from 68.9 in Single-Doc tasks to 55.3 in Multi (N-Gold) scenarios. This multi-gold setting consistently yields the lowest SAA scores across the board. 6 Conclusion We introduced CiteVQA, a benchmark designed to advance trustworthy document intelligence by requiring models to provide element-level visual citations alongside answers. By exposing these hidden hallucinations, CiteVQA establishes a rigorous standard for developing interpretable and reliable multimodal systems in high-stakes domains. 3. Contributions Our main contributions are threefold: • The CiteVQA Benchmark and Traceability Metrics: We introduce an evaluation framework that transitions Doc-VQA from answer-only scoring to joint evidence-answer verification, anchored by the Strict Attributed Accuracy (SAA) metric. • Scalable High-Fidelity Dataset Construction: We design an automated data generation pipeline that resolves the cost and consistency bottlenecks of granular visual annotation. • Discovery of the Attribution Hallucination
Phenomenon: Through a comprehensive audit of 20 leading MLLMs, we expose a critical vulnerability where models frequently output correct text while grounding it in entirely incorrect visual evidence. The paper concludes that CiteVQA provides the instrumentation needed to close the reliability gap that answer-only evaluations overlook. 3.4 Dataset Overview and Analysis As summarized in Table 2 and Figures 3-4, CiteVQA is a diverse benchmark comprising 711 documents across 7 macro-domains, with a realistic average length of 40.6 pages. The questions cover varied scenarios including single-doc (52.0%), multi-doc with one gold document (25.7%), and multi-doc with multiple gold documents (22.3%). Evidence is uniformly distributed across document positions and often spans multiple pages, demanding robust long-context aggregation. 4 Evaluation 4.1 Evaluation Metrics To evaluate evidence attribution, we introduce a novel set of metrics assessing both answer correctness and trustworthiness in grounding predictions on verifiable evidence. The key metrics include Recall (Rec.), Relevance (Rel.), Answer Correctness (Ans.), and Strict Attributed Accuracy (SAA), defined as SAA = 1(Ans.≥4∧(Rel.≥4∨Rec.≥0.6)). 4.3 Main Results Table 3 presents a comprehensive evaluation of state-of-the-art MLLMs on CiteVQA. The analysis reveals several critical insights into the current state of faithful evidence attribution. The Attribution Hallucination
Phenomenon A pervasive gap exists between answer accuracy (Ans.) and Strict Attributed Accuracy (SAA) across all tested models. Notably, while GPT-5.4 and Gemini-3-Flash achieve high answer scores (87.1 and 84.5), their SAA scores drop significantly to 59.0 and 65.4, respectively. Performance Disparity across Model Tiers There is a stark performance hierarchy among different model categories. Closed-source MLLMs dominate the benchmark, with Gemini-3.1-Pro-Preview leading at an Overall SAA of 76.0. In contrast, a significant cliff
exists for Open-source Models, where the strongest (Qwen3-VL-235B) achieves an SAA of only 22.5. 4.4 More Results of Experiments Widespread Deficiency in Coarse-grained Attribution A striking observation from Table 12 is that Page-level Recall (Page.) remains remarkably low for the vast majority of models. This indicates that the failure in evidence attribution is not merely a consequence of weak fine-grained grounding, but a more fundamental inability to navigate to the correct document page. Impact of Multi-document Complexity The challenge of attribution is significantly exacerbated as the environment shifts from Single-Doc to Multi-Doc (N-Gold) settings. In these high-density scenarios, even top-tier models exhibit a sharp performance collapse. 5. Analysis & Discussion 5.1 Fine-grained Results Question Type Results show a significant performance gap between question types. Models excel in Quantitative Reasoning (e.g., Gemini-3.1-Pro-Preview at 82.6) because numerical computations rely on objective logic and offer clear alignment between evidence and answers. In contrast, the newly introduced Multimodal Parsing task remains a major bottleneck; this category requires models to locate specific document elements based on descriptive cues and subsequently parse the content, leading to substantial difficulties in both precise evidence attribution and final answer generation. 5.2 Further Analysis of Evidence Attribution Beyond the initial identification of attribution fallacies, we seek to further explore the nuanced relationship between Attribution and Accuracy Synergy between Attribution and Accuracy Beyond serving as a metric for trustworthiness, faithful attribution appears to be positively correlated with the model’s reasoning success.
Improvements for AI systems
-
Evaluate models using Strict Attributed Accuracy (SAA) to ensure predictions are grounded in correct evidence, as
models frequently produce the right answer while citing the wrong region.
This directly addressesAttribution Hallucination
by rewarding models only whenthe answer and the cited region are both correct,
which is critical for high-stakes domains. -
Implement a scalable, automated annotation pipeline to generate high-fidelity, element-level bounding-box citations, as manual annotation is
prohibitively expensive and prone to inconsistencies.
This allows for the creation of a robust dataset comprising1,897 complex queries across 711 multi-page, multi-domain PDFs,
which improves the fidelity of document understanding. -
Develop an inference system that utilizes element-level citation rules, requiring evidence to be
at the element level: a complete paragraph, a complete table, a complete image, or a complete note.
This prevents models from selecting partial text and ensures that evidence isprecisely bounding box,
which is vital for verifiable reasoning. -
Introduce Traceability Metrics like Page-level Recall (Page.) to assess
coarse-grained localization ability,
as the paper notes thatPage-level Recall (Page.) remains remarkably low for the vast majority of models.
This helps diagnose whether failures are due to weak fine-grained grounding or a fundamental inability to navigate to the correct document page. -
Design a multi-agent framework, inspired by
DocDancer
andAgenticOCR,
that can performelement-level bijective function fmap
mapping between synthetic layouts and original files, ensuring thatevery byte of synthesized evidence is traceable to its original page.
This eliminates visual hallucinations by enforcing absolute fidelity in citation annotations.
Sources
- GPT-4 Technical Report
- Qwen3-VL Technical Report
- GAPS: A Clinically Grounded, Automated Benchmark for Evaluating AI Clinicians
- M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
- LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
- ColPali: Efficient Document Retrieval with Vision Language Models
- Know Or Not: a library for evaluating out-of-knowledge base robustness
- mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
- What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
- WebSailor: Navigating Super-human Reasoning for Web Agent
- ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios
- WebGPT: Browser-assisted question-answering with human feedback
- MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- Gemini: A Family of Highly Capable Multimodal Models
- Gemma: Open Models Based on Gemini Research and Technology
- MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale
- RARE: Retrieval-Augmented Reasoning Modeling
- AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering