FinInteract: Benchmarking Clarification and Intent Integration in Ambiguous Financial Question Answering
cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model agents increasingly answer financial questions by searching regulatory filings.
Terminology
Abstract
Large language model agents increasingly answer financial questions by searching regulatory filings. Such questions are often deceptively under-specified: Meta Platforms' "operating income" is 46.75B consolidated but 62.87B for the Family of Apps segment, and each reading is exactly verifiable against the filing. A capable agent should recognize the ambiguity and ask, rather than commit to a plausible but unintended reading. Existing financial benchmarks cannot measure this, because one gold answer per question cannot separate agents that resolve the ambiguity from those that guess the common reading, a blind spot we call the single-gold illusion. We release FinInteract, a bilingual (English/Chinese) benchmark of 173 instances that pairs each question with a default and an intended interpretation across a five-category ambiguity taxonomy, and grades whether an agent elicits the right clarification and then integrates it. Re-grading identical outputs against the default rather than the intended reading inflates GPT-4o's accuracy by 3.1 times, confirming the illusion. Beyond it, we find that models answer above 90% once the interpretation is supplied but at most 28.9% when they must elicit it themselves, that targeting is uneven across a taxonomy well powered for entity scope and metric definition and exploratory elsewhere, and that conditioning on the ambiguity category improves resolution at both inference and training time.
Sources
- Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks
- FinanceBench: A New Benchmark for Financial Question Answering
- ConvAI3: Generating Clarifying Questions for Open-Domain Dialogue Systems (ClariQ)
- CLAM: Selective Clarification for Ambiguous Questions with Generative Language Models
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models
- Knowing but Not Showing: LLMs Recognize Ambiguity but Rarely Ask Clarifying Questions
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection