V-FiLLM: Verified Financial LLM Reasoning Benchmark
Alicia Larsen, Victoire Laurent, Aulia Kharis Rakhmasari, Lara Turgut, Nino Antulov-Fantulin
ETH Zürich · Aisot Technologies Ltd
cs.AI, cs.CE, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-25
Comments: 10 pages, 6 tables, 2 figures, under review
Code: https://github.com/auliakharis/ML-in-Finance-and-Complex-System
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: V-FiLLM is a framework that generates financial reasoning benchmarks from executable computation trees grounded in real tables, yielding items whose answers are correct by construction.
Terminology
Summary
V-FiLLM is a framework that generates financial reasoning benchmarks from executable computation trees grounded in real tables, yielding items whose answers are correct by construction. Trees are evaluated symbolically to obtain ground truth and rendered into natural-language questions, removing any model from the labeling loop, so items can be generated at arbitrary scale without annotation cost and without inheriting a generator’s error rate. V-FiLLM exposes four independently controllable axes of difficulty including computation depth, expression breadth, financial concept complexity, and context size. By evaluating on open-source models, the authors find that accuracy falls up to 51% as reasoning depth increases, and up to 47% points under adversarial numerical perturbations, highlighting remaining challenges in robust financial reasoning over tables. They further show that lightweight LoRA fine-tuning on verified chain-of-thought traces improves accuracy from 81.1% to 85.6% on held-out problems and outperforms the base model by 5% points on FinQA, suggesting that targeted, low-cost adaptation is a promising direction for compositional reasoning in financial QA.
The benchmark generation pipeline maps synthetic financial spreadsheets to natural-language questions with executable answers. The first stage constructs two synthetic financial data sources: one mimicking a typical 10-Q filing and a regularized financial sheet with consistent column names, units, and temporal structure. The regularized sheet contains company-year observations for 15 synthetic companies from 2020 to 2025, with financial variables including revenue, costs, expenses, assets, liabilities, equity, cash, receivables, inventories, investments, current liabilities, capital expenditures, dividends, shares outstanding, stock price, and employees. Financial values are generated to preserve internal coherence while remaining synthetic. From these spreadsheets, typed financial atoms are constructed, each storing numerical value together with semantic metadata: financial concept, company, fiscal year, unit, and quantity type.
The next stage samples a symbolic expression template, represented as a typed binary tree. Leaves are financial atoms, while internal nodes specify an operation and metadata about their children. The operator set includes addition, subtraction, multiplication, ratios, growth rates, minimum, maximum, and average. The sampler is type-aware, only constructing expressions whose child outputs are compatible with the parent operation. Tree depth provides a direct proxy for the number of reasoning steps required: simple questions such as Provide value of company A's total assets in 2024
have depth 0, while more complex ones have greater depth. The pipeline also supports named derived financial concepts such as gross profit, operating income, pretax income, income tax expense, net income, current assets, long-term assets, and long-term liabilities, stored as special atoms with hidden executable formula trees. A rejection step prevents sampled templates from reconstructing canonical formulas of protected concepts without explicitly using the corresponding named concept atom.
After a valid template is sampled, the pipeline grounds it in concrete spreadsheet atoms using a binding environment that tracks constraints such as company, fiscal year, and financial concept. The bound expression is then rendered into a natural-language question through a bottom-up semantic procedure. Lightweight linguistic data augmentation is applied, introducing minor typographical errors, varying capitalization, and replacing standard phrasing with informal alternatives. The executable expression is evaluated directly to obtain the ground-truth answer, and for every generated example, the pipeline stores the natural-language question, numeric answer, full symbolic expression, template expression, requested and realized depths, probability of sampling derived concepts, and list of spreadsheet cells used as leaves.
The multi-turn extension converts expression trees into multi-turn dialogues, where each parenthesized sub-expression becomes one turn, ordered by computational dependency, with the final turn restating the original question in full. Per-turn targets are read directly from the corresponding nodes of the tree, requiring no additional annotation. The adversarial robustness extension perturbs input tables of existing items while leaving the question and answer unchanged, with four perturbation families: missing values, garbage values, OCR look-alikes, and cross-sheet contamination. Question-level perturbation appends irrelevant context such as rumors, market sentiment, and statements about other companies. The framework also allows controlling reasoning difficulty through depth, breadth via balanced trees, and value scaling via a multiplicative factor that enlarges operands without altering reasoning structure.
For LoRA fine-tuning, Low-Rank Adaptation is applied to attention projection layers (q proj, k proj, v proj, o proj) and feed-forward network projections (gate proj, up proj, down proj), with rank r=16, scaling factor α=32, and dropout rate 0.05. The fine-tuning recipe generates chain-of-thought traces for each training item and retains only those whose final answer matches the ground-truth value, yielding 608 verified examples from a pool of 688 (88% retention). The sweep of r, α, and dropout on the held-out benchmark (n=90) found that r=8, α=8, dropout 0.10 attains the highest accuracy (85.6%), tied at dropout 0.15, with accuracy most sensitive to the α/r ratio.
Benchmarking results show that model rankings remain stable across both evaluation settings, with Gemma-31B achieving the highest accuracy in both contexts (98.4% and 97.6%), closely followed by GPT-OSS-120B, Qwen3.7-Plus, and DeepSeek-v4-Flash. On simplified statements, Llama-3.3-70B and Qwen3.5-9B degrade to 80.8% and 55.2% respectively, whereas Gemma-31B and DeepSeek-v4-Flash exhibit high resilience. Multi-turn restructuring improves Qwen3.5-9B from 55.2% to 86.8% and Llama-3.3-70B from 80.8% to 86.8%. Accuracy stays steady up to 6 steps on simplified statements but then falls quickly, with Gemma-31B dropping from 85.0% at 6 steps to 55.0% at 8 steps, and DeepSeek-v4-Flash dropping from 84.0% to 26.0%. For 10-Q fillings, both models get almost perfect scores for 1 or 2 steps but drop to about 60–72% for 3 or 4 steps.
Under adversarial perturbations, all models perform much worse, with Gemma-31B's score dropping by 27.3 points on 10-Q Fillings and DeepSeek-v4-Flash dropping by 29.6 points on average across six noise types. Drops are even bigger on simplified statements (35.7 and 47.0 points). Adding useless extra information is the easiest test, lowering scores by less than 20 points on 10-Q Fillings. Changing units or scale is the most harmful test on 10-Q Fillings, crashing Gemma-31B to just 3.0% and DeepSeek-v4-Flash to 22.0%. When combining all noise together, scores fall to 26.0% and 17.0% on real documents. The LoRA-finetuned Qwen3.5-4B correctly answered 32/100 questions on FinQA, compared to 27/100 for the baseline.
The conclusion states that reasoning depth is the primary factor limiting model performance, with accuracy degrading substantially as the number of compositional reasoning steps increases. While LoRA fine-tuning on verified chain-of-thought traces improves performance over strong zero-shot baselines, considerable headroom remains, particularly on deeper reasoning tasks. OCR-style character corruption poses a greater challenge than missing or irrelevant data, highlighting the importance of robust numerical extraction in practical financial applications. The multi-turn benchmark provides fine-grained insight into intermediate reasoning failures not observable through end-to-end evaluation alone. Limitations include constrained compute resources, coverage of only six models, restriction to English-language documents, consideration of a single table-formatting convention, and fine-tuning experiments restricted to a relatively small evaluation set of 90 problems.
Improvements for AI systems
Improvements to AI systems:
-
Add a symbolic verification layer for chain-of-thought reasoning. The system generates CoT traces, evaluates the final answer against a symbolic executor, and discards traces with mismatched answers. This guarantees that training data is correct by construction, eliminating the propagation of reasoning errors from weaker generators.
-
Implement type-aware compositional reasoning constraints. The system enforces that every intermediate reasoning step operates on compatible semantic types (e.g., currency, ratio, growth rate). This prevents the model from combining incompatible quantities, reducing nonsensical intermediate computations in multi-step financial reasoning.
-
Introduce depth-controlled curriculum learning. The system trains on problems ordered by computation-tree depth (0 to 8+), progressively increasing reasoning complexity. This allows the model to master shallow reasoning before tackling deeper compositional chains, improving accuracy on high-depth tasks where current models drop by up to 51%.
-
Add adversarial robustness training via table perturbation. The system augments training data with four perturbation families—missing values, garbage values, OCR look-alikes (e.g., "0
→
O", "1→
l"), and cross-sheet contamination—while keeping questions and answers unchanged. This teaches the model to ignore irrelevant or corrupted table cells, addressing the 47-point accuracy drop observed under numerical perturbations. -
Implement multi-turn decomposition for complex queries. The system converts deep expression trees into multi-turn dialogues, where each sub-expression becomes a separate turn ordered by computational dependency. This breaks long reasoning chains into manageable steps, improving accuracy by up to 31.6 points (e.g., Qwen3.5-9B from 55.2% to 86.8%) on deep tasks.
-
Add unit and scale normalization layers. The system explicitly tracks units (e.g., millions vs. billions) and scale factors for every numerical value. It detects and corrects unit mismatches before computation, addressing the most harmful perturbation type that crashed accuracy to 3–22% when units were altered.
-
Incorporate OCR-error detection heuristics. The system flags character-level anomalies in table values (e.g., "l
instead of
1", "Oinstead of
0") and applies correction rules before reasoning. This targets the OCR-style corruption that proved more damaging than missing data or irrelevant context. -
Enable intermediate-step verification in multi-turn settings. The system checks the output of each turn against the corresponding sub-expression’s symbolic value. If a turn produces an incorrect intermediate result, the system halts and retries with alternative reasoning paths, preventing error accumulation across steps.
-
Implement adaptive difficulty selection during inference. The system estimates the reasoning depth of a query (via template matching or a lightweight classifier) and allocates more computational effort—such as higher sampling temperature or more verification passes—for deeper queries, mitigating the sharp accuracy drop observed at 8+ steps.
-
Add cross-document consistency checking. The system detects when a query references multiple tables (e.g., 10-Q filings and regularized sheets) and validates that values from different sources are consistent in units, fiscal years, and company identifiers before combining them, reducing cross-sheet contamination errors.
What the improved AI system can do:
-
Answer multi-step financial questions with up to 85% accuracy on 8-step reasoning tasks, compared to current models dropping to 26–55% at that depth.
-
Maintain accuracy within 10 points of baseline when tables contain OCR errors, missing values, or irrelevant appended context, instead of dropping 27–47 points.
-
Correctly handle unit and scale mismatches (e.g., millions vs. billions) without crashing, recovering from near-zero accuracy to >80% on unit-perturbed tables.
-
Decompose complex queries into verifiable sub-steps, allowing the system to pinpoint and correct the exact reasoning step that fails, rather than producing a wrong final answer.
-
Train on synthetic data with zero annotation cost, generating unlimited verified examples with controlled difficulty, enabling continuous improvement without human labeling.
-
Transfer learned compositional reasoning to real-world benchmarks like FinQA, outperforming baseline models by 5 percentage points after only lightweight fine-tuning on 608 verified examples.
-
Provide interpretable intermediate reasoning traces that can be audited step-by-step, making failures diagnosable and fixable in production financial QA systems.
Abstract
While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains comparatively less explored. We introduce V-FiLLM, a framework that generates financial reasoning benchmarks from executable computation trees grounded in real tables, yielding items whose answers are correct by construction. Trees are evaluated symbolically to obtain ground truth and rendered into natural-language questions, removing any model from the labeling loop, so items can be generated at arbitrary scale without annotation cost and without inheriting a generator's error rate. V-FiLLM exposes four independently controllable axes of difficulty including computation depth, expression breadth, financial concept complexity, and context size. By evaluating on open-source models, we find that accuracy falls up to 51% as reasoning depth increases, and up to 47% points under adversarial numerical perturbations, highlighting remaining challenges in robust financial reasoning over tables. We further show that lightweight LoRA fine-tuning on verified chain-of-thought traces improves accuracy from 81.1% to 85.6% on held-out problems and outperforms the base model by 5% points on FinQA (Chen et al., 2022a), s), suggesting that targeted, low-cost adaptation is a promising direction for compositional reasoning in financial QA.
Sources
- TabFact: A Large-scale Dataset for Table-based Fact Verification
- FinQA: A Dataset of Numerical Reasoning over Financial Data
- ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering
- DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs
- Faith and Fate: Limits of Transformers on Compositionality
- Parameter-Efficient Transfer Learning for NLP
- LoRA: Low-Rank Adaptation of Large Language Models
- FinanceBench: A New Benchmark for Financial Question Answering
- BizBench: A Quantitative Reasoning Benchmark for Business and Finance
- Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning
- GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- FinGPT: Instruction Tuning Benchmark for Open-Source Large Language Models in Financial Datasets
- GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
- MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual Data
- Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
- TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance
- DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection