Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting
cs.LG, cs.AI
Submitted: 2026-09-11
Updated: 2026-09-25
Comments: 35 pages, 9 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: In autoregressive forecasting, long prediction rollouts provide distant supervision, but backpropagation through time (BPTT) carries gradients from those losses through many autoregressive steps.
Terminology
Abstract
In autoregressive forecasting, long prediction rollouts provide distant supervision, but backpropagation through time (BPTT) carries gradients from those losses through many autoregressive steps. Repeated Jacobian products can make distant gradients dominate the update while amplifying predictable signal and unpredictable noise together; a large distant gradient therefore need not carry reliable learning signal. Motivated by this observation, we introduce Internal Dual-Wiener routing (Internal-DW), a principled backward-only intervention that preserves the full forward rollout and all horizon losses while reliability-weighting internal gradient routes. At each residual block, we derive bounded Wiener gains for the identity and nonlinear routes that balance preserving predictable learning signal against suppressing unpredictable variation, and estimate them from route-level gradient statistics and an explicit noise model. In a controlled system with known gradient signal-to-noise ratio (SNR), we show that distant gradients can grow even as their SNR falls, and that Internal-DW reduces held-out error in recovering predictable gradient signals and improves forecasting. On four history-dominated, weak-drive testbeds, Internal-DW reduces forecast error by 5.2%-13.8% relative to full BPTT, outperforms gradient clipping and Jacobian regularization on all four, and outperforms validation-selected truncated BPTT (TBPTT) on three. It also extends or preserves the fitted optimal training-horizon range across these four testbeds. Across the full benchmark suite, the current Internal-DW estimator has a clear applicability boundary: its benefit diminishes or reverses when usable history is limited or when the selected sampler fails to represent dominant drive-dependent variation. The results show that retaining long-horizon supervision does not require trusting every backward contribution equally.
Sources
- Training Deep Nets with Sublinear Memory Cost
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Dream to Control: Learning Behaviors by Latent Imagination
- Adam: A Method for Stochastic Optimization
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- Do Transformer World Models Give Better Policy Gradients?
- Jacobian-Adaptive Weighting for Stability: Enhancing Long-term Rollout of Neural Partial Differential Equation Solvers via Spatially-Adaptive Regularization
- Controlling Transient Amplification Improves Long-horizon Rollouts
- Synthetic Returns for Long-Term Credit Assignment
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- Unbiasing Truncated Backpropagation Through Time
- TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents
- Recurrent Neural Operators: Stable Long-Term PDE Prediction
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks