CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence

summary

Video file (mp4)

The gist

The gist The CiteVQA benchmark introduces an evaluation framework that requires models to return element-level bounding-box citations alongside each answer, evaluating both jointly.

In short

CiteVQA introduces a new benchmark that forces models to provide precise, element-level citations for every answer. The evaluation uses Strict Attributed Accuracy (SAA) to check if both the answer and its cited evidence are correct. Audits reveal 'Attribution Hallucination,' where models correctly answer but cite the wrong parts of a document, highlighting a major gap in multimodal reliability.

Key concepts

CiteVQA Benchmark
A new evaluation framework that requires AI models to return specific, element-level bounding-box citations alongside their answers. It tests whether the model can link its generated text directly to the exact visual source within a document.
Strict Attributed Accuracy (SAA)
The core metric used to judge model performance. A prediction only scores as correct if both the answer provided and the cited region of evidence are accurate. This ensures that models aren't just guessing; they must correctly identify and link verifiable facts.
Attribution Hallucination
A critical vulnerability discovered in models where they generate a correct answer but point to entirely incorrect visual evidence within the document. This phenomenon exposes a gap between linguistic understanding and accurate spatial grounding, showing models can be wrong about their sources even when their text is right.

Terminology used across episodes

This episode discusses

The paper

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence · Read on arXiv

Peking University

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence".

Jane: The gist The CiteVQA benchmark introduces an evaluation framework that requires models to return element-level bounding-box citations alongside each answer, evaluating both jointly.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, the CiteVQA paper sets up this whole evaluation framework where models have to return element-level bounding-box citations alongside their answers. They are using a metric called Strict Attributed Accuracy, or SAA, to judge both the answer and the cited region at the same time.

Jane: The core of their claim is that this joint evaluation exposes what they call Attribution Hallucination, which is when a model gives you the right answer but grounds it in completely incorrect visual evidence.

Lu: They built this benchmark around one thousand eight hundred ninety-seven questions spread across seven hundred eleven PDFs covering seven different domains and two languages, with documents averaging about forty point six pages each to keep it realistic <ref:2605.12882#pg1>.

Meng: The paper claims that by forcing models to provide these precise citations, they can establish a rigorous standard for measuring evidence fidelity in high-stakes areas like finance or law.

Lalam: It really matters because this isn't just about getting the answer right; it’s about knowing the specific visual source behind that claim, which is what makes an AI trustworthy for critical tasks.

Conclusion: Tom: So, looking at the CiteVQA benchmark, we see it’s a direct response to a weakness in how we evaluate document understanding before this work came out by Dongsheng Ma and his team.

Jane: The authors are essentially arguing that answer-only scoring is insufficient because it hides this specific failure mode where the model mixes up correct text with wrong visual evidence.

Lu: The implication here for the wider field is that we need to move beyond just scoring the final output and start demanding verifiable, element-level grounding from multimodal models.

Meng: For practical application, this means any system we build that reads documents needs this kind of accountability to ensure it’s not accidentally citing something irrelevant just because it sounds plausible.

Lalam: It settles a big question: can we build reliable document intelligence if we require every single claim to point directly to a specific location on the page?

Tom: Exactly, and the paper shows that while some models do well on answering things, their SAA scores are much lower than what's required for high-stakes work.

Jane: So, CiteVQA gives us the tools to test these systems under real conditions and show where they actually fall short in terms of accurate attribution.

More episodes

← Home