Into the ORBIT for Time Series: Training Regimes for Foundation Models

arXiv:2608.13262 · cs.LG, cs.AI · Submitted 2026-08-13 · Read on arXiv

Hongjie Xia, Yiding Liu, Yifan Hu, Peiyuan Liu, Zewei Dong

Ant International

cs.LG, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/ant-intl/Falcon-TST

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 95/100

The gist: This paper introduces ORBIT (Omni-Range Bootstrap Incremental Training), a training paradigm designed to explicitly control the effective pre-training distribution of heterogeneous time series

Terminology

Summary

This paper introduces ORBIT (Omni-Range Bootstrap Incremental Training), a training paradigm designed to explicitly control the effective pre-training distribution of heterogeneous time series corpora. The authors argue that while time series foundation models (TSFMs) have advanced primarily through architectural innovations—such as group attention, flow matching, serial-token prediction, and mixture-of-experts scaling—the training regimes governing data exposure over large-scale heterogeneous corpora remain comparatively under-explored. The paper states: the effective pre-training distribution is often poorly controlled along four coupled axes: cross-domain imbalance, frequency-dependent context requirements, variable prediction horizons, and missingness.

Under ORBIT, the authors train Falcon-2.0, a deliberately simple univariate encoder-only Transformer featuring missingness-aware triple-channel patch tokenization and direct multi-patch quantile prediction. The model is evaluated on GIFT-Eval and fev-bench, achieving strong zero-shot forecasting performance across diverse domains and frequencies.

The paper identifies four principal contributions:

  1. ORBIT paradigm: A unified training paradigm for heterogeneous temporal corpora that identifies effective pre-training distribution control as a critical but under-explored factor in TSFM development.

  2. Bootstrap Multi-Level Sampling and Omni-Range Incremental Training: Two complementary components that respectively control heterogeneous data exposure and enable single-stage training over diverse temporal contexts and forecasting horizons.

  3. Falcon-2.0 model: A simple encoder-only Transformer demonstrating that carefully designed training regimes can unlock strong forecasting capability without relying on excessive architectural complexity.

  4. Rank-Guided Cross-Depth Alignment: A training-time representation regularization method that transfers information from deeper Transformer layers to shallow layers without additional inference overhead.

ORBIT separates the construction of forecasting examples from their consumption during optimization through two complementary components:

This component controls data exposure hierarchically:

  • Corpus level: Prescribed domain-aware dataset weights are converted into an ordered global training stream using a low-discrepancy greedy blending rule. The paper notes: For each slot, the rule selects the dataset whose cumulative assigned count has the largest deficit relative to its target count at that point.

  • Dataset level: Bootstrap Stochastic Sampling constructs an offline sample index through four stochastic selections:

  • Level-1 Record Selection: Records shorter than 2P time steps are excluded; record index is sampled with equal probability from retained records.

  • Level-2 Target Variable Selection: A target-variable index is drawn with equal probability from available variables.

  • Level-3 Context Window Sampling: The split point is sampled uniformly from feasible positions, then context start offset is sampled uniformly from its feasible range.

  • Level-4 Prediction Horizon Sampling: Horizon length is sampled uniformly from feasible integer set P,..., min(T max, L rm - e m).

Each indexed example is represented by a five-tuple (r m, v m, s m, e m, p m) containing record, variable, context start, context end, and prediction horizon. This construction explicitly randomizes forecasting configurations while preserving reproducibility and efficient random access.

This component consumes variable-length examples throughout a single training run:

  • Context lengths and prediction horizons are sampled once during sample-index construction rather than resampled during optimization.

  • During batch assembly, contexts are left-padded and targets are right-padded to their respective mini-batch maxima, with attention and loss masks excluding unsupported positions.

  • The globally interleaved sample stream is consumed incrementally from memory-mapped storage.

  • "Short and long contexts, together with short and long prediction horizons, therefore coexist throughout optimization, avoiding the separate context-extension stages or horizon-specific training schedules adopted by several existing approaches."

Falcon-2.0 adopts a minimally specialized backbone following the univariate encoder formulation of Chronos-2, with several key components:

  • Missingness-Aware Reversible Instance Normalization: Instance statistics are computed only from observed values, with arcsinh transform applied. For entirely unobserved contexts, (µ, σ) = (0, 1) by convention.

  • Triple-Channel Patch Tokenization: Each patch is represented through temporal, value, and observation-indicator channels, producing a representation z ∈ R(3P). A shared residual SwiGLU projection maps each patch to the latent dimension.

  • Parallel Future Queries: The model constructs M future query patches for prediction, with temporal features defined as t fut j[κ] = (jP + κ)/C. Future target values are unavailable, so the value channel is zero and the indicator channel is set to one.

The encoder consists of D = 32 Pre-RMSNorm Transformer blocks with RoPE (base 10,000), output-gated self-attention, and SwiGLU feed-forward layers. The attention operation is summarized as:

O = softmax(Q attn K attn T / √d h + m) V attn ⊙ sigmoid(G)

The model uses bidirectional attention within the admissible set comprising observed context support, the REG token, and future query tokens.

A residual quantile head produces forecasts in the normalized arcsinh space with N q = 21 quantile levels: Q = 0.01, 0.05, 0.10,..., 0.90, 0.95, 0.99. The median channel (q = 0.5) provides the point forecast used by multi-stage inference.

This training-only objective uses late-layer representations as stop-gradient teachers to regularize shallow layers. The paper states: "We further show that sufficiently small alignment error bounds the perturbation between centered shallow and deep representations and, under an explicit spectral-separation condition, prevents the shallow representation from losing non-negligible modes present in the deep representation."

The token-wise cosine objective is:

L align = (1/n v) Σ ω b,u (1 - ⟨z sh b,u, z dp b,u ⟩)

where gradients are stopped through the deep representation. The default configuration uses alignment blocks (l sh, l dp) = (1, 31) with λ align = 10.0.

For horizons exceeding T max = 96 time steps, the model uses multi-stage prediction: prediction is parallel within each stage and autoregressive only across stages. Normalization is performed once, and the same statistics are held fixed at every stage. Only the median trajectory is fed back; all quantile trajectories are retained as outputs.

The default Falcon-2.0 configuration contains:

  • D = 32 encoder blocks

  • 585M trainable parameters

  • Latent dimension d = 1024

  • 16 attention heads with per-head dimension 64

  • FFN hidden size 4096

  • Patch size P = 16

  • Maximum context length C = 8192 time steps (512 context patches)

  • Maximum future patches M max = 6 (per-stage horizon T max = 96)

  • Maximum encoder tokens S max = 519

Falcon-2.0 is trained with:

  • 21-quantile pinball regression loss

  • AdamW optimizer (β1=0.9, β2=0.95)

  • Peak learning rate 6 × 10−5 with cosine annealing to 6 × 10−6

  • 1,000,000 training iterations

  • Weight decay 0.1, gradient clipping 1.0

  • Batch size 64 per GPU

  • BF16 mixed precision on NVIDIA B200-180GB clusters

  • Min/max prediction lengths: 16/96

The training corpus spans seven domains—Energy, Finance, Healthcare, Nature, Sales, Transport, and Cloud/IT—with rigorous separation from evaluation benchmarks following GIFT-Eval leakage prevention principles.

Falcon-2.0 achieves the strongest point-forecasting result among 29 evaluated pretrained models:

  • Lowest normalized MASE: 0.6684

  • Best mean MASE rank: 7.81

  • Improves over STRIDE + Timer-S1 (0.6744) by 0.9%

For probabilistic performance, Falcon-2.0 achieves the seventh-lowest CRPS (0.4843) with a mean CRPS rank of 9.62, though STRIDE + Chronos-2 retains an edge (0.4544; rank 6.84).

Falcon-2.0's aggregate normalized MASE of 0.6459 is within 0.3% of the top-performing TimesFM-2.5 (0.6438) and virtually tied with Chronos-2 (0.645), while establishing a superior mean MASE rank (5.15 vs. 5.63). Crucially, Falcon-2.0 achieves the best aggregate WQL (0.4842) among all models while successfully completing all 100 tasks.

The paper notes: "No single baseline outperforms Falcon-2.0 on both aggregate MASE and WQL simultaneously: TimesFM-2.5 yields slightly better point-forecasts but higher WQL, whereas Chronos-2 offers competitive mean ranks but inferior aggregate WQL."

On GIFT-Eval, Falcon-2.0's normalized MASE changes smoothly from 0.643 on short horizons to 0.683 and 0.725 on medium and long horizons, outperforming both Chronos-2 and Toto-2.0-2.5B at every horizon. The model leads on both univariate (0.641) and multivariate (0.705) tasks.

On fev-bench, the analysis reveals a critical architectural boundary: "On the 54 tasks without known-future covariates, Falcon-2.0 significantly outperforms Chronos-2 in both MASE (0.642 vs. 0.663) and WQL (0.467 vs. 0.476). However, on the 46 tasks with covariates, where Falcon-2.0's autoregressive interface does not ingest future features, this performance ordering reverses (MASE of 0.652 vs. 0.621; WQL of 0.509 vs. 0.498)."

On GIFT-Eval, Falcon-2.0 achieves the lowest normalized MASE in four of seven domains: Energy (0.769), Healthcare (0.531), Nature (0.650), and Transport (0.576). The largest margin occurs in Nature (10.2% reduction). Econ/Fin remains a weakness where Toto-2.0-2.5B dominates (0.739 vs. 0.785).

On fev-bench, Falcon-2.0 leads MASE in Cloud (0.566), Economy (0.623), Energy (0.642), and Mobility (0.665), while leading WQL in Economy (0.528), Healthcare (0.581), Mobility (0.553), and Nature (0.346).

Data scaling: From 100k to one million steps, MASE falls by 10.6% on GIFT-Eval and 12.1% on fev-bench, while CRPS on GIFT-Eval and WQL on fev-bench fall by 11.2% and 12.8% respectively.

Model scaling: Scaling from 75M to 585M parameters reduces GIFT-Eval MASE from 0.6839 to 0.6684 and CRPS from 0.4986 to 0.4843 (2.3% and 2.9% relative reductions). On fev-bench, MASE falls from 0.6745 to 0.6459 and WQL from 0.5083 to 0.4842 (4.2% and 4.7% reductions).

Architecture ablations: Parallel Patch Prediction has the largest effect—removing it increases MASE and CRPS on GIFT-Eval by 8.0% and 8.7%, and MASE and WQL on fev-bench by 4.1% and 8.0%. Triple-Channel Patch Tokenization and residual SwiGLU patch projection provide smaller but consistent improvements (0.7–1.9% and 0.4–1.0% respectively).

Sampling ablations: Bootstrap Stochastic Sampling reduces MASE and CRPS on GIFT-Eval by 11.7% and 13.6% respectively, and MASE and WQL on fev-bench by 5.4% and 6.5% compared to sliding-window enumeration with the same context/horizon sampling. Joint sampling of context lengths and prediction horizons is the only configuration achieving the lowest error on all four metrics.

The paper concludes: These results highlight training-distribution design as a key factor in building scalable and generalizable time series foundation models. The authors demonstrate that explicitly controlling source exposure and interleaving context and horizon ranges during training translates into a model competitive across distinct forecasting regimes, while the residual performance gaps (particularly on covariate-rich tasks) define concrete directions for future refinement.

Improvements for AI systems

Based on the paper, here are specific improvements I can implement in AI systems:

  1. Implement Bootstrap Multi-Level Sampling for any time series model training, replacing naive sliding-window enumeration. This hierarchical sampling (record → variable → context → horizon) reduces error by 5.4–13.6% across metrics by explicitly randomizing forecasting configurations while maintaining reproducibility.

  2. Adopt Omni-Range Incremental Training to eliminate separate context-extension stages or horizon-specific schedules. By interleaving short/long contexts and horizons throughout optimization, the model learns robust representations across all regimes simultaneously, avoiding catastrophic forgetting of short-horizon patterns when scaling to long horizons.

  3. Apply Rank-Guided Cross-Depth Alignment as a training-only regularizer. This transfers information from deeper layers to shallow layers via stop-gradient cosine alignment, improving representation quality without any inference overhead—particularly valuable for resource-constrained deployment where shallower models must approximate deeper ones.

  4. Integrate Missingness-Aware Triple-Channel Patch Tokenization (temporal + value + observation-indicator channels). This enables the model to explicitly distinguish observed vs. unobserved values, improving robustness to missing data—a common real-world issue that degrades standard models.

  5. Use Parallel Future Queries with Multi-Stage Autoregression for long-horizon prediction. This balances parallelism (within-stage) and sequential dependency (across stages), reducing error accumulation while maintaining computational efficiency compared to fully autoregressive approaches.

  6. Adopt Quantile Forecasting Heads with 21 levels for probabilistic output. This provides calibrated uncertainty estimates without requiring multiple model runs, enabling risk-aware decision-making in downstream applications.

  • Achieve state-of-the-art zero-shot forecasting across diverse domains (energy, finance, healthcare, nature, sales, transport, cloud/IT) without fine-tuning, with 0.9% better point forecasts than the previous best model on GIFT-Eval.

  • Handle heterogeneous data distributions by explicitly controlling cross-domain imbalance, frequency-dependent context needs, and variable prediction horizons—making it suitable for mixed-source industrial data streams.

  • Maintain performance under missing data through missingness-aware normalization and tokenization, reducing the need for imputation preprocessing.

  • Provide calibrated probabilistic forecasts (CRPS, WQL) alongside point estimates, enabling uncertainty-aware planning, inventory management, and risk assessment.

  • Scale efficiently with predictable improvements: 10.6–12.8% error reduction when doubling training steps, and 2.3–4.7% improvement when scaling model size from 75M to 585M parameters.

  • Operate without architectural complexity—the simple encoder-only Transformer with these training regimes outperforms more complex models, reducing inference cost and deployment complexity.

  • Handle both short and long horizons (16 to 96+ steps) in a single model, eliminating the need for horizon-specific models or multi-stage training pipelines.

Abstract

Time series foundation models (TSFMs) have advanced primarily through architectural innovation, while training regimes for large-scale heterogeneous corpora remain under-explored. As a result, pre-training distributions are often poorly controlled with respect to domain imbalance, context requirements, prediction horizons, and missingness. We introduce ORBIT (Omni-Range Bootstrap Incremental Training), a training paradigm that makes this distribution explicit and controllable. ORBIT combines Bootstrap Multi-Level Sampling, which controls dataset exposure and samples records, target variables, context windows, and prediction horizons, with Omni-Range Incremental Training, which varies context lengths and prediction horizons throughout a single training stage. Under ORBIT, we train Falcon-2.0, a simple univariate encoder-only Transformer with missingness-aware triple-channel patch tokenization and parallel patch prediction. We further introduce Rank-Guided Cross-Depth Alignment, a training objective that uses late-layer representations as stop-gradient teachers for shallow layers without additional inference cost. Evaluations on GIFT-Eval and fev-bench demonstrate strong zero-shot forecasting performance across diverse domains and frequencies.

Sources

Related papers