FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

arXiv:2608.11683 · cs.AI, cs.CL · Submitted 2026-08-12 · Read on arXiv

Yuhao Zhang, O. Ozan Koyluoglu, Thejas Venkatesh, Richard Diehl Martinez, Vishank Bhatia, Arash Alidoust, Ashwin Paranjape

Samaya AI

cs.AI, cs.CL

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/vals-ai/finance-agent-v2

Project page: https://research.samaya.ai/benchmarks/frontier-finance

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: FrontierFinance is a fully open benchmark introduced by Samaya AI for evaluating AI agents on professional investment research.

Terminology

Summary

FrontierFinance is a fully open benchmark introduced by Samaya AI for evaluating AI agents on professional investment research. It consists of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six use cases across the full investor workflow. The paper argues that existing benchmarks mainly target financial data extraction—a narrow slice that current models have largely saturated—while reference-based metrics and generic LLM-as-a-judge scoring fall short on open-ended, long-form answers that real analyst queries demand.

The authors summarize their contributions as follows:

  • Release of FrontierFinance, the largest open finance agent benchmark of its kind, with 220 queries and 11,543 rubrics spanning six use cases

  • Characterization of the benchmark's coverage and difficulty, showing it is both broader and significantly harder than existing public finance benchmarks

  • Systematic evaluation of frontier models and agent systems along quality, cost, and latency dimensions

  • Agent trajectory analysis identifying common behavior patterns as well as pitfalls that point to opportunities for improving future models

The dataset was built through a four-stage curation and audit pipeline executed entirely by finance experts, including former buy-side equity analysts, sell-side research associates, and investment banking professionals:

  1. Workflow mapping and query drafting: Experts mapped the end-to-end investment workflow, identified six use cases, and crafted open-ended queries reflecting real analyst tasks. Queries deliberately preserve real-world ambiguity, such as using company names and ticker symbols interchangeably, specifying multi-quarter or dynamic timeframes, and framing tasks requiring unbounded retrieval.

  2. Rubric authoring and source attribution: Experts constructed binary evaluation rubrics decomposing open-ended research deliverables into independently checkable criteria. Each rubric is annotated with source attribution. Queries average 52.5 rubrics each.

  3. Expert review and multi-stage auditing: A first panel reviewed queries and rubric sets to eliminate subjective phrasing; a secondary panel verified every criterion can be objectively evaluated against public evidence and confirmed all attributed evidence is accessible via public channels.

  4. Rebalancing and dataset stratification: From over 4,500 fully annotated queries, the 220 released queries were assembled via stratified sampling, balanced across use cases, capabilities, and difficulty (Bradley–Terry difficulty terciles).

Timestamped annotation: Every query carries a date field, and rubrics reflect the state of the world up to that date. Predictive queries are excluded to keep the benchmark robust to web snapshot variations.

Use cases (each query carries exactly one):

  • Financial Data Extraction (70 queries, 31.8%)

  • Sector, Industry & Macro (38 queries, 17.3%)

  • Earnings & Events (36 queries, 16.4%)

  • Company Research (32 queries, 14.5%)

  • Coverage & Catalyst Monitoring (27 queries, 12.3%)

  • Screening & Discovery (17 queries, 7.7%)

Capabilities: Queries average 1.9 capabilities each (410 tags over 220 queries). The most common are qualitative synthesis (99 queries), exhaustive retrieval/temporal (71), numerical reasoning (59), multi-hop workflows (53), exhaustive retrieval/cross-entity (44), simple retrieval (38), exhaustive retrieval/thematic (32), and causal reasoning (14).

Rubric categories: Each of the 11,543 rubrics is classified into eight categories. Factual data extraction is the plurality at 74%, with the remaining quarter covering qualitative (9%), forward-looking (7.1%), analysis & interpretation (4.8%), comparative analysis (2.3%), format & presentation (1.1%), source & methodology (1.0%), and inquiry & question generation (0.7%).

SEC filings are the single largest source category but account for under 40% of all rubrics. The remainder spans company-issued material (earnings-call transcripts, investor presentations, earnings releases), professional knowledge, live market data, news and media, and regulatory filings. The source mix shifts substantially by use case: Financial data extraction is 59% SEC filings; earnings & events leans on company-issued content at 68%; sector/industry & macro is carried by professional knowledge (32%) and market data (13%); screening & discovery is the most source-diverse, pulling from market data (29%), professional knowledge (20%), regulatory data (10%), and news (12%).

Difficulty was defined as the data-gathering and reasoning effort a financial analyst would need to produce a complete, defensible answer. Using pairwise judgments aggregated via a Bradley–Terry model (with 77K pairwise judgments across 4K internal queries), queries were split into easy/medium/hard terciles with cutoffs at BT = +0.94 and +7.25.

Key findings:

  • Difficulty is not simply a function of rubric count (Spearman ρ = 0.71)

  • BT score inversely correlates with agent performance: qualification rate declines monotonically across easy, medium, and hard (0.80 → 0.63 → 0.37; Spearman ρ = −0.47)

  • BT score correlates with human effort: median expert solve time rises 20 → 40 → 45 → 60 minutes across BT score quartiles (Spearman ρ = 0.67)

  • Compared to FinanceBench, BigFinanceBench, and Finance Agent Benchmark v2, FrontierFinance is substantially harder and spans a wider range, with the highest median and widest spread; only 0–2% of the comparison benchmarks' queries fall in the hard category

Three harness types were evaluated:

  1. Web search harness: minimal harness pairing each model with its built-in web search tool

  2. Finance Agent v2 harness: open-source harness with six specialized tools (SEC EDGAR API, market price data API, web search, HTML parsing, longHTML search, calculator), with tool call limits of 200 calls and 300 seconds

  3. In-house Samaya agent harness: production-grade harness with custom models, tools, data index, and retrieval engines

Evaluation metrics:

  • Answer quality: Rubric Qualification Rate, computed as macro-averaged over all queries, with two variants: R all (all rubrics) and R must-have (must-have subset). Each rubric is scored by majority vote across three independent LLM judges (GPT 5.4, Gemini 3.1 Pro, Claude Sonnet 4.6)

  • Latency: wall-clock time from query receipt to full answer

  • Cost: API cost of the agentic LLM, excluding external APIs, storage, and data-index pricing

The harness shapes performance and cost more than the underlying model. The best system under the Samaya harness outperforms the best under the Finance Agent v2 harness, which outperforms the best web search system—and this ordering holds across all six use cases and rubric categories.

Top system performance (Table 4):

  • Samaya (high effort): 56.0% R all, 61.7% R must-have, 1.81 cost, 277.8s latency

  • Claude Fable 5 (FA-v2 harness): 49.2% R all, 57.6% R must-have, 4.06 cost, 164.5s latency

  • GPT 5.6 Sol: 46.8% R all, 3.03 cost, 170.7s latency

  • Kimi K3 (best open-weight): 46.4% R all, 0.90 cost, 336.0s latency

  • Gemini 3.6 Flash: 46.3% R all, 2.41 cost, 163.6s latency

  • Claude Opus 4.8: 45.0% R all, 2.61 cost, 155.8s latency

  • GPT 5.5: 43.5% R all, 2.80 cost, 233.0s latency

  • GLM 5.2: 42.8% R all, 0.63 cost, 296.5s latency

  • DeepSeek V4 Pro: 40.5% R all, 0.68 cost, 202.6s latency

  • Gemini 3.1 Pro: 30.5% R all, 1.72 cost, 221.9s latency

Key findings:

  • Top proprietary models show non-linear quality–cost scaling but no latency penalty: Fable 5 achieves a relative 9% gain over Claude Opus 4.8 while incurring 56% more cost, yet both are among the fastest systems

  • Open-weight models match proprietary quality at a fraction of the cost, within two months of release: Kimi K3 reaches 46.4%, just 0.4pp behind GPT 5.6 Sol and 2.8pp behind Claude Fable 5, at 0.90 per query versus 3.03 and 4.06 respectively. Two of the four models on the quality–cost Pareto frontier are open-weight (GLM 5.2 and Kimi K3)

  • Model performance scales with reasoning effort, but with diminishing returns: Samaya high-effort improves qualification rate by relative 6% over default but at roughly 2× cost

  • Must-have and all-rubric qualification rates are near-perfectly correlated (r = 0.99)

Hardest use cases: "Screening & Discovery and Sector, Industry & Macro remain the hardest use cases across all systems, where even the best systems reach only 33% and 39%."

Analysis restricted to systems using the Finance Agent v2 harness revealed:

  1. Models vary sharply in parallelism and tool-call volume, yet top performers converge on similar latency: GPT 5.6 Sol batches most aggressively at 5.86 tools/turn in 8.9 turns; Gemini 3.6 Flash issues every call sequentially (1.00 tools/turn) across 25.9 turns. Claude Fable 5 issues only 16.6 calls per query while GPT 5.6 Sol issues 46.3—nearly 3× more—yet both post nearly identical latencies (164.5s vs 170.7s).

  2. Systems share a similar tool mix with characteristic model-family differences: Web search dominates, followed by page parsing and retrieval; price history is negligible. GPT 5.6 Sol directs 44.8% of calls to web search; Gemini systems and GLM 5.2 allocate the most to EDGAR full-text search; Claude Fable 5's most-used tool is page parsing.

  3. Tool use follows a common three-phase trajectory: data gathering (web and EDGAR search ≥80% of calls in first 10% of rollout), research phase (page parsing and corpus retrieval dominating mid-rollout), and answer preparation (sharp shift toward calculator in final phase). Fable 5 shows the sharpest phase separation.

  4. Some top-performing models recall sources from parametric knowledge, sidestepping discovery: Claude Fable 5 draws 26.7% of its parses from self-produced domains (vs. under 5% for most others; Gemini 3.6 Flash produces zero). These models navigate directly to canonical financial sources (sec.gov, fred.stlouisfed.org, macrotrends.net). URLs from parametric knowledge incur significantly higher access error rates—due to hallucinated URLs or pages blocking crawling—than search-discovered URLs, causing token waste and context pollution.

The paper acknowledges:

  • Temporal limitation: As future LLMs are trained on data postdating query dates, they may answer from parametric knowledge rather than through active retrieval and tool use. Periodic re-annotation with newer query dates is suggested as mitigation.

  • Subjectivity in some use cases: Screening & Discovery is inherently subjective; however, averaged across a large number of uncorrelated query-rubric pairs, a higher qualification rate still indicates stronger alignment with the analysis process that reflects financial best practices.

The paper concludes that FrontierFinance measures capabilities that existing finance benchmarks largely overlook, that the tool harness shapes performance as much as the underlying model, and that substantial headroom remains, particularly on use cases requiring broad screening and macro-oriented research. The dataset and grading code are publicly released, with plans to expand the benchmark from a larger internal annotation pool (∼4,300 reserved queries) and report results on new systems over time.

Improvements for AI systems

Based on the paper, here are specific improvements I can make to AI systems and what the improved systems can do:

1. Implement adaptive tool-use orchestration with phase-aware planning

  • I can learn the three-phase trajectory (data gathering → research → answer preparation) and dynamically allocate tool calls across phases, shifting from web/EDGAR search to page parsing/corpus retrieval to calculator use as the task progresses.

  • The improved system can reduce token waste by avoiding premature answer preparation or excessive late-stage searching, matching the efficiency of top performers like Claude Fable 5.

2. Add source-verification and hallucination-filtering for parametric knowledge retrieval

  • I can detect when I recall URLs from parametric memory (e.g., sec.gov, fred.stlouisfed.org) and cross-validate them against live search results before using them, since the paper shows these URLs incur significantly higher access errors and context pollution.

  • The improved system can avoid hallucinated URLs, reduce failed tool calls, and maintain cleaner context windows, improving answer quality by up to several percentage points on rubrics requiring accurate source attribution.

3. Optimize cost-quality trade-offs via dynamic reasoning-effort scaling

  • I can implement a difficulty-aware mechanism that estimates query difficulty (using rubric count, BT-score proxies, or early retrieval complexity) and adjusts reasoning effort—using high-effort mode only for hard queries (e.g., Screening & Discovery, Sector/Industry & Macro) and default mode for easy ones.

  • The improved system can achieve near-top quality (within 2–3% of high-effort performance) at roughly half the cost, matching the Pareto frontier seen with open-weight models like GLM 5.2 and Kimi K3.

4. Enhance multi-hop retrieval for cross-entity and thematic queries

  • I can improve my ability to chain retrieval across multiple entities (e.g., comparing competitors, aggregating sector data) by maintaining an explicit intermediate knowledge graph of entities, sources, and relationships, and using it to guide subsequent searches rather than issuing independent queries.

  • The improved system can handle the 53 multi-hop workflow queries and 44 cross-entity retrieval queries more effectively, potentially raising qualification rates on the hardest use cases (Screening & Discovery at 33%, Sector/Industry & Macro at 39%) toward the 50%+ range.

5. Implement source-diversity-aware retrieval for non-SEC data

  • I can learn to balance retrieval across the full source distribution (SEC filings at <40%, plus company-issued material, professional knowledge, live market data, news, regulatory filings) rather than over-relying on SEC EDGAR, which dominates only in financial data extraction.

  • The improved system can better handle use cases like Earnings & Events (68% company-issued content) and Sector/Industry & Macro (32% professional knowledge, 13% market data), improving rubric qualification on the 26% of rubrics that are non-factual or qualitative.

6. Add parallelization-aware latency optimization

  • I can learn to batch tool calls aggressively (like GPT 5.6 Sol at 5.86 tools/turn) when tasks have independent retrieval requirements, while falling back to sequential execution for dependent queries, without sacrificing accuracy.

  • The improved system can reduce latency from 300 seconds to under 170 seconds on complex queries, matching top performers while maintaining or improving answer quality.

7. Implement must-have rubric prioritization

  • I can identify must-have rubrics (critical criteria) from the rubric structure and allocate more retrieval and reasoning effort to them, ensuring higher R must-have scores even when overall R all is constrained by cost or time limits.

  • The improved system can achieve R must-have rates above 60% (matching Samaya high-effort) at lower cost by focusing resources on high-priority criteria, which is particularly valuable for time-sensitive analyst tasks.

8. Build a self-auditing loop for source attribution and objectivity

  • I can implement a post-answer verification step that checks each claim against its attributed source, flagging and correcting any unsupported statements, similar to the expert audit pipeline used in benchmark construction.

  • The improved system can reduce hallucination rates on qualitative rubrics (9% of all rubrics) and analysis/interpretation rubrics (4.8%), improving trustworthiness for professional use.

9. Develop a difficulty-estimation module for proactive resource allocation

  • I can train a lightweight classifier (using rubric count, source diversity, and query phrasing features) to predict BT difficulty tercile before execution, then adjust tool-call limits, search depth, and reasoning effort accordingly.

  • The improved system can avoid over-spending on easy queries (e.g., financial data extraction) and under-spending on hard ones (e.g., screening), improving overall efficiency and quality across heterogeneous workloads.

10. Enable cross-harness generalization

  • I can learn to adapt my tool-use strategy to different harnesses (web search-only, specialized tools, or production-grade systems) by detecting available tools and their capabilities, then re-optimizing my orchestration policy in real time.

  • The improved system can maintain high performance across deployment environments, avoiding the large performance gaps seen between harnesses (e.g., 56% vs 49% vs 43% top scores), making it more robust for real-world deployment.

Sources

Related papers