Fewer yet critical: Reducing Redundant Token Dependencies for Transformer-based Time Series Forecasting
cs.AI, cs.LG
Submitted: 2025-03-10
Updated: 2026-09-08
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Time series forecasting (TSF) is important in real-world applications.
Terminology
Abstract
Time series forecasting (TSF) is important in real-world applications. Recently, Transformer-based methods have achieved strong performance by modeling token dependencies through attention mechanisms. However, existing methods are usually trained mainly with prediction error losses, which may cause models to exploit both critical and redundant token dependencies. Such redundant dependencies can introduce irrelevant information and weaken generalization. To address this issue, we propose a simple yet effective token dependency selection strategy. Specifically, by jointly introducing the attention entropy constraint and prediction error constraint, the model can identify fewer but more critical inter-token dependencies and perform forecasting based on them, thus avoiding the interference of redundant dependencies. The proposed method can be easily extended to various Transformer-based TSF models. Experiments on multiple TSF datasets demonstrate its effectiveness.
Sources
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting
- When can transformers reason with abstract symbols?
- Long-term Forecasting with TiDE: Time-series Dense Encoder
- SimMTM: A Simple Pre-Training Framework for Masked Time-Series Modeling
- iTransformer: Inverted Transformers Are Effective for Time Series Forecasting
- Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification
- Deep Transformer Models for Time Series Forecasting: The Influenza Prevalence Case
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection