Clean Scores, Buried Evidence, and Confident Wrong: A Receipt-Based Audit of Frontier Agentic QA
cs.IR, cs.AI, cs.CL, q-fin.GN
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 28 pages, 9 figures. Frozen evidence archive: https://doi.org/10.5281/zenodo.22310532
License: http://creativecommons.org/licenses/by/4.0/
The gist: Frontier models score well on shallow document/chart reading tasks.
Terminology
Abstract
Frontier models score well on shallow document/chart reading tasks. In a controlled data-room audit, moving evidence into buried conditions reduced accuracy, increased forced declarations, increased tool calls, and increased cost per correct answer. Confidence and benchmark calibration did not fully capture wrong answers; a documented production incident shows fabricated structural claims can be mixed with accurate numeric tables. Agentic evaluations need claim-level receipts (statement-level provenance, not answer-level scores), condition-aware scoring, and human-adversarial verification - an auditing discipline, not a leaderboard. The setting we measure is financial due diligence; the setting we are building toward next is defense staff work, where the same buried-evidence shape appears. In both, the model is not a party to the consequences; the person who signs is. In plain terms: in the documented cases we examine, agents can pair accurate numbers with confident fabricated explanations, and the burden of proof must therefore move from the model to the evidence trail.
Sources
- Holistic Evaluation of Language Models
- The Embedder's Dilemma: LLMs Are Better, but at What Cost?
- GAIA: a benchmark for General AI Assistants
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
- MuSiQue: Multihop Questions via Single-hop Question Composition
- Lost in the Middle: How Language Models Use Long Contexts
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Why Language Models Hallucinate
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- FinQA: A Dataset of Numerical Reasoning over Financial Data
- On Calibration of Modern Neural Networks
- Language Models (Mostly) Know What They Know
- Teaching Models to Express Their Uncertainty in Words
- Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- Selective Question Answering under Domain Shift
- Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG