CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA
cs.CL, cs.AI
Submitted: 2026-08-13
Updated: 2026-09-07
Comments: Accepted at FinNLP 2026 @ EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Existing defenses against hallucination in retrieval-augmented and multi-agent pipelines remain partial: evidence is trusted despite modality disagreement, debate verifies an aggregate report rather
Terminology
Abstract
Existing defenses against hallucination in retrieval-augmented and multi-agent pipelines remain partial: evidence is trusted despite modality disagreement, debate verifies an aggregate report rather than individual claims, and such verification occurs only after drafting, leaving inter-agent errors undetected until the final text. To close this gap, we present CLAIR-Fin, a nine-agent framework that decomposes each question into atomic claims maintained in a typed Financial Claim Ledger. Each claim is resolved through Asymmetric Evidence Authority, which conditions evidence trust on claim type rather than treating all modalities as equally reliable; Chain-of-Custody Verification, which checks grounding at the hand-off between drafting and adversarial review rather than only at the pipeline's exit; an Adaptive Rebuttal Cycle, which routes contested claims through adversarial debate whose depth scales with what that debate finds; and a terminal entailment audit paired with a continuous Hallucination Risk Index that distinguishes claims that passed scrutiny from claims never contested. We evaluate CLAIR-Fin on BB-FinQA-X, a 500-question cross-modal financial evaluation set built from Bangladesh Bank Annual Report material, stratified by query type, format, and difficulty. Relative to a single-pass retrieval-augmented generation baseline, it raises faithfulness (0.780 to 0.889) while abstaining on 5.4% of questions when evidence is insufficient rather than forcing an unsupported response, and it exceeds stronger retrieval-strategy baselines such as HyDE and Graph-RAG on faithfulness (at most 0.874).
Sources
- FinGround: Detecting and Grounding Financial Hallucinations via Atomic Claim Verification
- A Comprehensive Survey of Retrieval-Augmented Generation (RAG): Evolution, Current Landscape and Future Directions
- MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
- FinanceBench: A New Benchmark for Financial Question Answering
- Chain-of-Verification Reduces Hallucination in Large Language Models
- Claim Verification in the Age of Large Language Models: A Survey
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- Toward Reliable Evaluation of LLM-Based Financial Multi-Agent Systems: Taxonomy, Coordination Primacy, and Cost Awareness
- FAMMA: A Benchmark for Financial Domain Multilingual Multimodal Question Answering
- Evaluating Large Language Models on Financial Report Summarization: An Empirical Study
- XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
- FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain
- FRED: Financial Retrieval-Enhanced Detection and Editing of Hallucinations in Language Models
- ChartHal: A Fine-grained Framework Evaluating Hallucination of Large Vision Language Models in Chart Understanding
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering