Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting
Kiran Madhusudhanan, Christian Klötergens, Lars Schmidt-Thieme, Vijaya Krishna Yalavarthi
University of Hildesheim
cs.LG, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting proposes TORF, a framework that decouples mean forecasting from uncertainty estimation to address the trade-off
Terminology
Summary
Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting proposes TORF, a framework that decouples mean forecasting from uncertainty estimation to address the trade-off between distributional flexibility and accurate mean prediction in probabilistic time series forecasting. The paper states: "Traditional parametric methods, such as Mean Variance Estimation (MVE), can suffer from degraded point accuracy when trained under joint Negative Log-Likelihood (NLL) objectives, while modern-flexible generative models, including Normalizing Flows and Diffusion Models, typically rely on costly Monte Carlo sampling and may yield suboptimal mean estimates."
The proposed TORF framework works in two stages: "In the first stage, a pre-trained deterministic model is used to produce an accurate mean prediction. In the second stage, a Restricted Normalizing Flow, with strictly odd functions learns flexible residual distributions around the point forecast, guaranteeing mean preservation from the first stage without sampling."
The paper identifies a fundamental problem with single-stage approaches: "Direct NLL training Figure 1b fails to recover the true mean in high-noise regions because the loss function heavily prioritizes low-noise regimes (Seitzer et al. 2022), whereas a two-stage approach that first fits the mean and then the variance recovers both Figure 1c. Furthermore, flexible density models
inherit poor mean accuracy for two distinct reasons. First, like MVE, they optimize a distributional objective (e.g., the ELBO or a log-likelihood loss) rather than a mean-focused loss, so the predictive mean is only an implicit by-product of fitting the full density and is never directly supervised. Second, unlike MVE, they lack an analytical mean altogether: recovering it requires expensive Monte Carlo sampling."
The key contributions listed in the paper are:
-
We show that a SimpleTM based two-stage Gaussian baseline (MVE-2S) already matches or beats K 2 VAE on 11/12 distributional-accuracy comparisons (Table 6).
-
We introduce TORF, which extends this decoupled framework with an odd-constrained residual flow (ROSS), modeling flexible, non-Gaussian residuals while provably preserving the exact Stage-1 mean, without sampling.
-
We design a lightweight CNN architecture that efficiently maps time-series contexts to ROSS's parameters.
-
TORF with SimpleTM (Chen et al. 2025) as 1st stage model sets a new state of the art, winning 7/9 CRPS and 9/9 NMAE on long-horizon forecasting and 6/8 CRPS and 5/8 NMAE on short-horizon forecasting.
The methodology section describes the two-stage framework: "Stage 1: Mean Estimation. Any point-predictor fθ1 trained independently under an MSE objective to produce the conditional mean estimate Ŷ = fθ1(X). TORF is model-agnostic at this stage i.e., any point-predictor can serve as fθ1. Stage 2: Residual Density Estimation. Let ε = Y−fθ1(X) denote the residual. Since Y 7→ ε is a translation, its Jacobian is the identity and pε(ε X) = pY(fθ1(X)+ε X). With θ1 frozen, a normalizing flow pθ2 is trained to model p(ε X) directly, constrained by construction to satisfy Epθ2[ε X] = 0."
The paper makes two assumptions on the residual distribution: 1. Conditional Independence. p(ε X) = ∏ p(εh,c X) over c,h. 2. Symmetry. p(ε X) = p(−ε X).
These assumptions trade a small amount of distributional generality for the exact mean-preservation guarantee at the core of our method.
The core architectural components include:
-
Residual Odd Splines (ROSS): "We prevent this by restricting the spline to the class of odd functions... Let S be a monotonic rational-quadratic spline defined over [0, B], defaulting to the identity outside. Rather than mirroring bin parameters across both domains, we process only the magnitude of the input through S and restore the sign (sgn(·)) afterward."
-
Linear Scaling Layer (LSL): "a translation term shifts the residual distribution, breaking odd symmetry and decoupling the predictive mean of p(Y X) from fθ1(X). LSL therefore restricts to scaling only, which is an odd function and preserves Epθ2[ε X] = 0."
-
SplineNet:
a lightweight 1D convolutional network operating over the time axis of F: Φ = Conv1D(ReLU(Conv1D(F)))... This reduces parameter cost to O(k · Chid · Nb) for kernel size k ≪ H, independent of the forecast horizon.
-
ScaleNet: "The scale matrix S ∈ RH×C >0 is predicted from F via a bounded MLP applied independently at each channel."
The paper proves that Since both components (LSL, ROSS) are strictly odd functions, their composition is also odd, analytically enforcing Epθ2[ε X] = 0 throughout.
When K=0, TORF reduces to a conditional Gaussian model, exactly recovering the MVE-2S baseline.
Experimental results show: "In the long-term setting, TORF wins on CRPS for 7 of 9 datasets and on NMAE for all 9 datasets. The largest gains reach +35.4% on CRPS (ETTh2-L) and +25.1% on NMAE (ETTm2-L). In the short-term setting, TORF wins on CRPS for 6 of 8 datasets and on NMAE for 5 of 8 datasets. The largest gains reach +30.3% on CRPS (ETTm2-S) and +28.1% on NMAE (ETTm2-S)."
The ablation study shows that TORF-1S underperforms the two-stage variants. Removing LSL in TORF (w/o LSL) substantially weakens TORF across all horizons and datasets.
Also, "all odd-constrained variants (Odd-TORF) are provably mean-preserving and inherit an identical Stage-1 mean, whereas the non-odd w RealNVP and w Spline can change the mean, visibly degrading NMAE (up to 0.629-0.630 vs. 0.583 on Solar-24), exactly the failure mode the odd constraint rules out."
The paper concludes: "We presented TORF, a two-stage probabilistic framework for TSF that pairs a frozen, model-agnostic point forecaster with a constrained odd flow whose predictive mean exactly matches the Stage 1 forecast, while a lightweight Conv1D SplineNet keeps it efficient enough to scale to long horizons. Across 17 real-world tasks, TORF achieves state-of-the-art results, with gains of up to +35.4% CRPS and +28.1% NMAE over the strongest baseline."
Limitations noted: "TORF assumes independent, symmetric residuals, which is sufficient here, but limiting when residuals are asymmetric or cross-dimensionally dependent. The Mixture extension relaxes symmetry at a higher parameter cost for only marginal gains, and joint time/channel dependence remains unaddressed."
Improvements for AI systems
Based on this paper, here are the specific improvements you can make to AI systems:
1. Decouple mean estimation from uncertainty modeling in any regression/prediction system.
Instead of training a single model with a joint likelihood objective, split the pipeline: first train a deterministic point predictor under MSE, then freeze it and train a separate density model on the residuals. This prevents the common failure where high-noise regions dominate the loss and degrade point accuracy. The improved system will produce both accurate point forecasts and calibrated uncertainty without sacrificing either.
2. Enforce exact mean preservation in generative density models via odd-function constraints.
When using normalizing flows or similar flexible density estimators, restrict the transformation to strictly odd functions (e.g., odd splines, odd scaling layers). This guarantees that the predictive mean of the residual distribution is exactly zero by construction, eliminating the need for costly Monte Carlo sampling to recover the mean. The improved system will have an analytical, exact mean that matches the point predictor, while still capturing non-Gaussian, asymmetric-free residual shapes.
3. Use a two-stage Gaussian baseline as a strong, cheap baseline for probabilistic forecasting.
Before deploying complex generative models, implement a simple two-stage approach: fit a point predictor, then fit a Gaussian distribution on residuals (MVE-2S). This baseline already outperforms many sophisticated variational or flow-based single-stage models on distributional accuracy. The improved system will have a fast, robust fallback that often beats more complex methods, saving compute and training time.
4. Replace Monte Carlo mean estimation with closed-form mean from the flow’s base distribution.
In any flow-based uncertainty model, design the flow so that the base distribution has zero mean and the transformation is odd. Then the predictive mean is exactly the point forecast, no sampling needed. This reduces inference cost and variance, making the system suitable for real-time or low-latency applications where sampling is prohibitive.
5. Add a lightweight convolutional context encoder for long-horizon forecasting.
Use a small 1D CNN (e.g., two Conv1D layers with ReLU) to map the time-series context to the flow’s parameters, rather than a heavy recurrent or attention-based encoder. This reduces parameter count and computational cost, scaling efficiently to long horizons without loss of accuracy. The improved system will handle long sequences faster and with less memory.
6. Introduce a linear scaling layer (LSL) that only scales, never shifts, to preserve mean.
In residual density models, avoid translation terms that break odd symmetry. Use only a scaling operation (which is odd) to adjust variance. This keeps the mean exactly at the point forecast while allowing flexibility in spread. The improved system will maintain guaranteed mean accuracy even when the residual distribution has heavy tails or heteroscedasticity.
7. Apply the two-stage framework to any existing point predictor (model-agnostic).
Because the first stage is frozen and any point predictor can be used, you can upgrade an existing deterministic model (e.g., a Transformer, LSTM, or CNN) into a probabilistic forecaster without retraining it. The improved system will add uncertainty quantification to legacy models with minimal changes, preserving their already-tuned point accuracy.
8. Use the ablation findings to avoid common pitfalls in flow-based uncertainty models.
Specifically, avoid single-stage NLL training for flows (which degrades mean accuracy), avoid non-odd transformations (which shift the mean and hurt NMAE), and avoid removing the scaling layer (which weakens performance). The improved system will follow these design rules to ensure both distributional flexibility and point accuracy.
What the improved AI system can do:
-
Produce accurate point forecasts and calibrated probabilistic predictions simultaneously, even in high-noise or heteroscedastic settings.
-
Generate non-Gaussian predictive distributions (e.g., skewed, heavy-tailed) without sacrificing the exactness of the mean.
-
Run inference without Monte Carlo sampling, making it fast enough for real-time forecasting.
-
Scale to long horizons with a lightweight architecture that uses less compute and memory.
-
Upgrade any existing deterministic forecaster into a probabilistic one with minimal effort and no loss in point accuracy.
-
Outperform state-of-the-art single-stage generative models on both CRPS (distributional accuracy) and NMAE (point accuracy) across diverse real-world time-series datasets.
Abstract
Probabilistic forecasting plays an essential role in risk-sensitive decision-making, particularly in long-horizon settings. However, existing approaches often face a fundamental trade-off between distributional flexibility and accurate mean prediction. Traditional parametric methods, such as Mean Variance Estimation (MVE), can suffer from degraded point accuracy when trained under joint Negative Log-Likelihood (NLL) objectives, while modern-flexible generative models, including Normalizing Flows and Diffusion Models, typically rely on costly Monte Carlo sampling and may yield suboptimal mean estimates. To address this limitation, we propose Two-stage Odd Residual Flows (TORF), a framework that decouples mean forecasting from uncertainty estimation. In the first stage, a pre-trained deterministic model is used to produce an accurate mean prediction. In the second stage, a Restricted Normalizing Flow, with strictly odd functions learns flexible residual distributions around the point forecast, guaranteeing mean preservation from the first stage without sampling. Experiments show that TORF achieves state-of-the-art deterministic accuracy (NMAE) while providing strong density estimation performance (CRPS) on short and long-horizon forecasting.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks