t 0: A Time-Series Foundation Model for Forecasting with Context
cs.LG, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 39 pages, 16 figures, 13 tables
Code: https://github.com/theforecastingcompany/tfc-t0
License: http://creativecommons.org/licenses/by/4.0/
The gist: We present t 0, a family of open-weights foundation models for forecasting with multivariate context.
Terminology
Abstract
We present t 0, a family of open-weights foundation models for forecasting with multivariate context. We release its first two members: t0-alpha and t0-beta, respectively 102M and 256M parameters. Both condition their forecasts on target history, past covariates, and known-future covariates, without task-specific retraining. Their transformer layers alternate attention along time and across variates. They produce probabilistic forecasts through quantile predictions. Pretraining combines curated public data with synthetic generator families constructed to contain covariate-to-target dependencies. On GIFT-Eval, t0-alpha reaches an aggregate CRPS of 0.4941, and t0-beta a CRPS of 0.4738 and a MASE of 0.6865, third on both and within 4.0% of the best zero-shot TSFM. On fev-bench they score 42.2 and 46.7 in skill, the latter third again and 2.0 points behind the leader. We analyze t0-alpha in depth. Known-future covariates raise its skill by 6.3 percentage points across 30 tasks. The report also examines its calibration, its rollout strategy on long horizons, and its robustness to missing data. On the Victoria electricity-demand benchmark, t0-beta is among the most accurate models with a context of nearly a year. In an independent Macrocosm evaluation of hourly ERCOT prices over 29 months, both cut the MAE of the lagged-price baseline by 38%.
Sources
- GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation
- Chronos: Learning the Language of Time Series
- Chronos-2: From Univariate to Universal Forecasting
- Zero-Shot Time Series Forecasting with Covariates via In-Context Learning
- TiRex: Zero-Shot Forecasting Across Long and Short Horizons with Enhanced In-Context Learning
- xLSTM: Extended Long Short-Term Memory
- Toto: Time Series Optimized Transformer for Observability
- A decoder-only foundation model for time-series forecasting
- In-Context Fine-Tuning for Time-Series Foundation Models
- Scaling Vision Transformers to 22 Billion Parameters
- Fewer Truncations Improve Language Modeling
- Metadata Matters for Time Series: Informative Forecasting with Transformers
- ForecastPFN: Synthetically-Trained Zero-Shot Forecasting
- Time to Embed: Unlocking Foundation Models for Time Series with Channel Descriptions
- MQTransformer: Multi-Horizon Forecasts with Context Dependent and Feedback-Aware Attention
- Query-Key Normalization for Transformers
- From Tables to Time: Extending TabPFN-v2 to Time Series Forecasting
- Toto 2.0: Time Series Forecasting Enters the Scaling Era
- Efficient Sequence Packing without Cross-contamination: Accelerating Large Language Models without Impacting Performance
- TS-ICL: A Flexible Time-Indexed Foundation Model for Time Series via In-Context Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks