Long-Horizon Forecasting of Complete Financial Statements with Forma
Travis L. Johnson, Jiannan Jiang, Soumyabrata Chaudhuri, Yihao Chen, Lauren Falvey, Donal O'Cofaigh
University of Texas at Austin
cs.LG, q-fin.CP
Submitted: 2026-08-11
Updated: 2026-08-13
Comments: 46 pages, 2 figures, 3 tables. Benchmark: https://github.com/forma-lab-mccombs/proforma-20q. Model and weights: https://github.com/forma-lab-mccombs/forma-release
Code: https://github.com/forma-lab-mccombs/proforma-20q
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: The paper introduces ProForma-20Q, a reproducible benchmark for forecasting 78 statement line items 1–20 quarters ahead for anonymized firms, and Forma, a transformer-based architecture tailored to
Terminology
Summary
The paper introduces ProForma-20Q, a reproducible benchmark for forecasting 78 statement line items 1–20 quarters ahead for anonymized firms, and Forma, a transformer-based architecture tailored to this task. The authors state: To our knowledge, no prior work jointly forecasts complete financial statements beyond one year, yet in a discounted-cash-flow valuation most firm value sits past that window.
The paper argues that the economically relevant object is therefore a joint forecast of the whole statement at horizons where value lives.
Specialist beats generalist: Specialist training beats generalist scale when forecasting financial statements.
Forma, a 0.9M-parameter transformer, beats every competitor we field: classical machine learning, chained gradient boosting, a zero-shot time-series foundation model, and frontier large language models.
Widening lead with horizon: Its lead widens with horizon, where valuation needs accuracy most.
Forma outperforms every model at every horizon beyond h=2, reaching at least 3.2 percentage points (pp) of R2 by h=20.
The LLMs' performance is generally poor: the best frontier model underperforms all but the simplest purpose-trained model, and its R2 deficit relative to Forma grows from 5.5pp at h=1 to 15.9pp at h=20.
Probabilistic forecasts: Forma produces Gaussian predictive intervals [that] never under-cover at any horizon.
Its forecasts nearly satisfy accounting identities; exact coherence is recoverable at no statistically significant accuracy cost.
Scenario analysis: Its tuple interface supports scenario analysis without retraining, and we show that pinning future revenue paths sharpens the rest of the statement.
The task is defined as: each example is indexed by firm f and forecast origin quarter t; horizon h refers to absolute quarter t+h. The historical input is a set of tuples: Sf,t = (h, id, x) ∶ h = −11, …, 0, id ∈ D, id reported at t+h.
The forecasting task is to learn x̂ = g(Sf,t, cf; h, id)
for h=1,…,20 and id∈D.
Data: The sample is quarterly U.S. filings (Compustat), excluding financial firms (SIC 6000–6999).
Values are deflated by scale (origin-quarter total liabilities + shareholders’ equity, which nearly always equals total assets), asinh-transformed, and standardized per (item, quarter).
Splits are temporal: train 1971–2001, validation 2002–2009, test 2010–2024,
with targets purged at boundaries. The panel spans 32,851 firms and 1,173,598 firm-quarters (609,269 train, 211,367 validation, 352,962 test).
Evaluation: Models are ranked by out-of-sample R2 for predicting changes.
The paper notes: In change space, R2 of 20–40% at multi-year horizons is strong, and level-space intuitions of 90% or higher do not apply.
Significance is measured via Diebold and Mariano [11] tests, accounting for dependence across cells in the same quarter by collapsing to calendar-quarter means.
Tuple-set representation: Each observed historical tuple (h, id, x) ∈ Sf,t is a token, and each requested future pair (h, id) is a query token whose value is hidden from the encoder.
The initial representation is: z(0)h,id = Eacct(id) + Ehorizon(h) + Evalue(x, m),
where Eacct is a learned account embedding, Ehorizon a fixed sinusoidal encoding, and Evalue uses a learned projection for observed values or a mask vector for hidden ones.
Key design choices:
-
An unreported historical item contributes no tuple and hence no token
— missingness is native, avoiding imputation. -
Two context tokens complete the input: industry and origin-quarter scale deflator.
-
The encoder has
4 layers, dmodel = 128, and 4 attention heads (≈0.9M parameters).
Training objective: Masked prediction with two masking variations:
-
Identity-aware grouped masking:
When masking touches a complete identity group we therefore mask at least two of its members
to avoid teaching constraint algebra. -
Pinned-future masking:
Half the training examples mask the entire future (pure forecasting); the other half reveal ≈5% of reported future tuple values as inputs
— enabling scenario analysis.
Output: For each requested pair (h, id), we map u = Concat[z(L)h,id, Ehorizon(h)] to a location and a heteroskedastic scale.
The primary head is a heteroskedastic Gaussian (mu, sigma), trained with the beta-NLL loss (beta = 0.5).
Five seeds are trained and treated as an equal-weight mixture.
Full sample (Panel A of Table 2): Forma explains 28.9% of the cross-sectional variance of realized changes, ahead of the RF (27.2%), penalized regression (25.8%), and both FFNN baselines (25.3% and 24.7%). Every gap is significant at the 1% level.
Horizon dynamics: "The RF slightly outperforms Forma one quarter ahead (R2 39.8% vs. 39.0%; DM t=+12.3), but the models are at parity at h=2, and from h=3 Forma is ahead with significance growing through h=20 (0.225 vs. 0.193; t reaching −18.6)."
Mechanism: "Fundamentals mean-revert at conditional, item-specific rates. The seasonal random walk misses reversion entirely (R2 =−4.1%). The pooled fade/AR(1) baseline captures unconditional reversion and explains 18.3% of the variation, over half of Forma’s total. What separates the models is the conditional component—reversion speeds that depend on the rest of the statement."
Architecture vs. capacity: "The larger FFNN (≈4.2M parameters) underperforms its linear sibling (24.7% R2 vs. 25.3%; ≈3.1M parameters), and both trail the ≈0.9M-parameter transformer by 3.6–4.2 percentage points. Architecture, not capacity, drives Forma’s edge."
Chained GBM: The L2 GBM variant achieves an R2 of 18.5%, compared with 24.7% for Forma on their common sample.
On the absolute-error track, the original L1 specification of Geertsema et al. [15] posts an MAE of 0.400, compared with 0.364 for the MAE-targeting Laplace Forma.
LLMs: Forma achieves 29.9% on the LLM sample, compared to 18.6% for the top-performing LLM. In fact, every purpose-trained model other than fade/AR(1) outperforms every LLM.
The comparison is conservatively biased toward the generalists as test-period financial statement realizations sit in their pretraining data.
LLM MAEs (0.362–0.368) are only slightly higher than that of the Laplace Forma (0.348)
— consistent with conditional medians, these forecasts are competitive in absolute error but underperform under the squared-error criterion essential for valuation.
Probabilistic quality (Panel C): Among the mixtures, the exact five-seed NLL and closed-form CRPS rank Forma’s Laplace variant clearly first and its Gaussian variant second.
The five-seed Gaussian mixture is essentially calibrated in total variance (z̄2=0.96).
Its central intervals never under-cover at any horizon—pooled coverage is 72.6/89.6/93.7/95.8% at nominal 50/80/90/95%.
The PIT is center-heavy, so the intervals are conservative rather than sharp.
Across 124.7M enforced identity instances, the median absolute violation of Forma’s raw-dollar statements is 3.7% of the identity’s gross scale.
The paper notes: part of the violation is a property of the estimand rather than prediction error
because means do not commute with the nonlinear asinh transform.
Reconciliation: "The variance-weighted projection drives violations to numerical zero at no statistically significant squared-error cost (R2 drops by 3.8 percentage points, quarter-clustered DM t=−1.4) while yielding a small but significant MAE improvement (t=+4.8). The equal-weight variant
is catastrophic (R2 falls to −5.51, MAE rises to 0.635) —
a dollar adjustment spread uniformly across accounts is negligible for total assets but enormous relative to small line items."
Pinning the true revenue path lowers pooled MAE on the remaining items from 0.409 to 0.383 and raises change-space R2 from 30.5% to 34.8%.
The gain is small one quarter out (0.9pp) and widens with horizon to 7.4pp by h=20.
"Income-statement and balance-sheet items gain equally in absolute error (ΔMAE −0.032 each). Balance-sheet items gain most in R2 (+8.9pp vs. +5.6pp). Cash-flow items barely move (ΔR2 +0.7pp), reflecting their weak link to revenue."
Long-horizon evaluation conditions on realized survivors: comparisons are fair (identical cells for all models), but absolute skill levels describe only the survivors.
Claims are scoped to quarterly U.S. filings, these 78 accounting items, and horizons of 1–20 quarters.
The paper measures forecast quality only; valuation applications require additional inferences about discount rates and terminal values.
It does not incorporate stock market or analyst data as additional features, and only consider off-the-shelf generalists rather than fine-tuned ones.
The paper's four contributions are:
-
Task and protocol: ProForma-20Q,
a turnkey protocol for complete financial-statement forecasting, scored in change space on common samples.
-
Architecture-to-problem fit: Forma's tuple-set transformer with
identity-aware masking forces the model to learn economics instead of accounting algebra; pinned-future masking trains it to condition on chosen realizations for scenario analysis.
-
Probabilistic statement forecasts:
Predictive densities over the statement via a heteroskedastic head and five-seed mixture
withexact ex-post reconciliation at statistically insignificant accuracy cost.
-
The specialist-wins-and-widens result:
Forma outperforms all competitors, including LLMs, and that its advantage widens with horizon.
Improvements for AI systems
Based on the paper, here are specific improvements to AI systems and their resulting capabilities:
Improvement: Replace the bigger is better
paradigm with a compact, domain-structured transformer (0.9M parameters) that uses tuple-set tokenization (account ID + horizon + value) and native missingness handling (no imputation). Train with identity-aware grouped masking to prevent the model from learning accounting algebra shortcuts, forcing it to learn economic relationships.
Capability: A system that outperforms frontier LLMs by 15.9pp R2 at 20-quarter horizons, with accuracy widening over time—critical for long-horizon valuation where most firm value resides. It achieves this with 300× fewer parameters than LLMs, enabling deployment on edge devices or in real-time financial pipelines.
Improvement: Architect the model to explicitly learn item-specific, condition-dependent mean-reversion speeds (rather than pooled/unconditional reversion). Use the finding that reversion speeds depend on the rest of the statement
to build attention mechanisms that capture cross-statement dependencies.
Improvement: Train with a dual-masking scheme: 50% pure forecasting, 50% with 5% of future values revealed as inputs. This creates a tuple interface
where users can pin specific future revenue paths without retraining.
Improvement: Implement a heteroskedastic Gaussian head with β-NLL loss (β=0.5) and five-seed mixture ensembles. Calibrate to ensure central intervals never under-cover at any horizon (observed: 72.6/89.6/93.7/95.8% at nominal 50/80/90/95%).
Improvement: Post-process forecasts with a variance-weighted projection onto accounting identities (assets = liabilities + equity), rather than equal-weight reconciliation. The paper shows equal-weight is catastrophic (R2 falls to −5.51), while variance-weighted costs only 3.8pp R2 (statistically insignificant) and improves MAE.
Improvement: Adopt the paper's evaluation methodology: score models on out-of-sample R2 for changes (not levels), use Diebold-Mariano tests with calendar-quarter clustering, and report on common samples with purged temporal splits.
Improvement: Acknowledge and handle the limitation that long-horizon evaluation conditions on realized survivors. Build systems that explicitly model firm survival probability and condition forecasts on both survival and non-survival paths.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks