CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
summary
The gist
The gist The CiteVQA benchmark introduces an evaluation framework that requires models to return element-level bounding-box citations alongside each answer, evaluating both jointly.
In short
CiteVQA introduces a new benchmark that forces models to provide precise, element-level citations for every answer. The evaluation uses Strict Attributed Accuracy (SAA) to check if both the answer and its cited evidence are correct. Audits reveal 'Attribution Hallucination,' where models correctly answer but cite the wrong parts of a document, highlighting a major gap in multimodal reliability.
Key concepts
- CiteVQA Benchmark
- A new evaluation framework that requires AI models to return specific, element-level bounding-box citations alongside their answers. It tests whether the model can link its generated text directly to the exact visual source within a document.
- Strict Attributed Accuracy (SAA)
- The core metric used to judge model performance. A prediction only scores as correct if both the answer provided and the cited region of evidence are accurate. This ensures that models aren't just guessing; they must correctly identify and link verifiable facts.
- Attribution Hallucination
- A critical vulnerability discovered in models where they generate a correct answer but point to entirely incorrect visual evidence within the document. This phenomenon exposes a gap between linguistic understanding and accurate spatial grounding, showing models can be wrong about their sources even when their text is right.
Terminology used across episodes
This episode discusses
- CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence · Paper Radio
- GPT-4 Technical Report
- Qwen3-VL Technical Report
- GAPS: A Clinically Grounded, Automated Benchmark for Evaluating AI Clinicians
- M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
- LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
- ColPali: Efficient Document Retrieval with Vision Language Models
- Know Or Not: a library for evaluating out-of-knowledge base robustness
- mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
- What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
- WebSailor: Navigating Super-human Reasoning for Web Agent
- ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios
- WebGPT: Browser-assisted question-answering with human feedback
- MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- Gemini: A Family of Highly Capable Multimodal Models
- Gemma: Open Models Based on Gemini Research and Technology
- MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale
- RARE: Retrieval-Augmented Reasoning Modeling
- AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
The paper
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence · Read on arXiv
Peking University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence".
Jane: The gist The CiteVQA benchmark introduces an evaluation framework that requires models to return element-level bounding-box citations alongside each answer, evaluating both jointly.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, the CiteVQA paper sets up this whole evaluation framework where models have to return element-level bounding-box citations alongside their answers. They are using a metric called Strict Attributed Accuracy, or SAA, to judge both the answer and the cited region at the same time.
Jane: The core of their claim is that this joint evaluation exposes what they call Attribution Hallucination, which is when a model gives you the right answer but grounds it in completely incorrect visual evidence.
Lu: They built this benchmark around one thousand eight hundred ninety-seven questions spread across seven hundred eleven PDFs covering seven different domains and two languages, with documents averaging about forty point six pages each to keep it realistic <ref:2605.12882#pg1>.
Meng: The paper claims that by forcing models to provide these precise citations, they can establish a rigorous standard for measuring evidence fidelity in high-stakes areas like finance or law.
Lalam: It really matters because this isn't just about getting the answer right; it’s about knowing the specific visual source behind that claim, which is what makes an AI trustworthy for critical tasks.
Conclusion: Tom: So, looking at the CiteVQA benchmark, we see it’s a direct response to a weakness in how we evaluate document understanding before this work came out by Dongsheng Ma and his team.
Jane: The authors are essentially arguing that answer-only scoring is insufficient because it hides this specific failure mode where the model mixes up correct text with wrong visual evidence.
Lu: The implication here for the wider field is that we need to move beyond just scoring the final output and start demanding verifiable, element-level grounding from multimodal models.
Meng: For practical application, this means any system we build that reads documents needs this kind of accountability to ensure it’s not accidentally citing something irrelevant just because it sounds plausible.
Lalam: It settles a big question: can we build reliable document intelligence if we require every single claim to point directly to a specific location on the page?
Tom: Exactly, and the paper shows that while some models do well on answering things, their SAA scores are much lower than what's required for high-stakes work.
Jane: So, CiteVQA gives us the tools to test these systems under real conditions and show where they actually fall short in terms of accurate attribution.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language