Temporal horizons in forecasting: a performance-learnability trade-off
cs.LG, nlin.CD
Submitted: 2025-06-04
Updated: 2026-09-09
Comments: 38 pages, 12 figures Permanent link with reviews: https://openreview.net/forum?id=BeudQIxT1R
Journal ref: Transactions on Machine Learning Research (TMLR), October 2025
License: http://creativecommons.org/licenses/by/4.0/
The gist: When training autoregressive models to forecast dynamical systems, a critical question arises: how far into the future should the model be trained to predict for optimal performance? In this work, we
Terminology
Abstract
When training autoregressive models to forecast dynamical systems, a critical question arises: how far into the future should the model be trained to predict for optimal performance? In this work, we address this question by analyzing the relationship between the geometry of the loss landscape and the training time horizon. Using dynamical systems theory, we prove that loss minima for long horizons generalize well to short-term forecasts, whereas minima found on short horizons result in worse long-term predictions. However, we also prove that the loss landscape becomes rougher as the training horizon grows, making long-horizon training inherently challenging. We validate our theory through numerical experiments and discuss practical implications for selecting training horizons. Our results provide a principled foundation for hyperparameter optimization in autoregressive forecasting models.
Sources
- Preferential Temporal Difference Learning
- AntisymmetricRNN: A Dynamical System View on Recurrent Neural Networks
- Learning to Plan for Language Modeling from Unlabeled Data
- Lipschitz Recurrent Neural Networks
- Were RNNs All We Needed?
- Better & Faster Large Language Models via Multi-token Prediction
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Improved memory in recurrent neural networks with sequential non-normal dynamics
- On the difficulty of training Recurrent Neural Networks
- Stepping on the Edge: Curvature Aware Learning Rate Tuners
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks