Evaluating Time Series Foundation Models for Electricity Price Forecasting: Contamination Risk, Distributional Shifts, and Covariate Dependence
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Evaluating Time Series Foundation Models for Electricity Price Forecasting".
Tom: Time series foundation models (TSFMs) have shown strong zero-shot forecasting performance, but their generalization in covariate-driven, nonstationary settings is underexplored.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, to wrap up what we talked about earlier, this paper by Pan and Ezzat sets up a rigorous two-dataset benchmarking framework specifically for electricity price forecasting to stop contamination risks. They focus on the core thesis that while time series foundation models show strong zero-shot performance, their ability to generalize in nonstationary settings is still an area needing serious exploration.
Jane: Precisely, Tom. The authors propose this framework because electricity price forecasting presents a tough testbed due to the complex temporal dependencies and distributional shifts inherent in the data. Their main claim is that they examine four key areas: point and probabilistic performance, tail behavior under extreme conditions, robustness to those nonstationary shifts caused by price spikes, and comparisons against domain-specific forecasting methods.
Lu: The methodology involves formalizing the electricity price forecasting task within a day-ahead market setup where the objective is to estimate the predictive distribution P (Yt+δt:t+δt+H−1Ft), which means they are looking at predicting a whole range of future prices based on all available information at time t <ref:2607.02623#pg0>.
Meng: That formal setup sounds like it covers everything from historical data to real-time weather and generation forecasts, which is exactly what we deal with in the energy sector. I’m interested in how they handle the input information Fj, since that’s where the covariate dependence comes into play.
Lalam: It seems like their focus on "covariate support" is central; they aren't just asking if the model works generally, but whether it relies on specific external factors to achieve that performance. This distinction is vital when building practical systems.
Tom: Right, and the results section starts by showing that TSFMs are competitive against many general-purpose baselines on a standardized benchmark called GEFCom2014-P, achieving competitive zero-shot performance without needing task-specific training or feature engineering.
Jane: But they immediately pivot to showing that the performance really hinges on whether those models have covariate support; they find that TSFMs generally outperform statistical and deep learning baselines when the models incorporate relevant features.
Lu: The paper also goes into a detailed analysis of probabilistic forecasting scores using Average Quantile Loss, which is a specific metric for assessing how well the model predicts uncertainty across different price levels. They highlight that certain variants, like Chronos two w <ref:2607.02623#pg1>. and TabPFN-TS w., achieve the best quantile-based performance among the TSFMs they tested.
Meng: That's interesting because accuracy isn't everything in forecasting; knowing how uncertain we are about a price is often just as important for risk management decisions, which is what those quantile scores measure.
Lalam: And then there’s the analysis of tail behavior, where they look at how models perform under extreme price conditions. They found that median performance alone doesn't tell the whole story because several models show substantial degradation in the lower tails when faced with these extreme events, while TSFMs produce more balanced forecasts across both extremes.
Tom: That’s a key finding there—that distributional robustness is important—but they also found that simple ensembles, like averaging the best TSFM and a domain-specific model called CING-LEAR, achieved the strongest overall performance. It suggests combining pre-trained knowledge with domain expertise is effective.
Jane: So, to summarize this segment, the paper establishes a way to fairly evaluate these models by looking at their point accuracy, probabilistic scores, tail behavior during spikes, and how they react when given specific covariate support compared to other methods.
Conclusion: Tom: Looking at the full picture of this paper, "Evaluating Time Series Foundation Models for Electricity Price Forecasting: Contamination Risk, Distributional Shifts, and Covariate Dependence" by Pan and Ezzat really highlights a crucial reality in applying large AI models to real-world energy tasks. The authors are essentially warning us that zero-shot performance isn't enough if we don't control for the data source or the specific context of the electricity market.
Jane: Exactly, Tom. The paper emphasizes that while these foundation models can achieve impressive initial forecasting scores, their true utility depends heavily on whether they have been trained with the right contextual information relevant to electricity prices, and how they handle those sudden shifts in market conditions.
Lu: From a research perspective, the implication is that we need to move beyond just scaling up model size and start focusing intensely on designing architectures that are inherently sensitive to the structural biases of real-world time series data, like price spikes or weather correlations.
Meng: For practical implementation, this means we can’t just deploy a generic model and hope for the best; we need a process to verify if the model is actually leveraging the necessary inputs before it goes live in our energy infrastructure.
Lalam: I think this research points toward an evolution in how we build forecasting tools; instead of relying purely on massive pretraining, we should be designing systems where domain structure and task-specific inductive biases are baked into the AI from the start.
Tom: And that ties back to their conclusion: accurate and robust performance in electricity price forecasting is strongly tied to domain structure and those task-specific inductive biases they mentioned earlier. It’s a strong statement about the necessary ingredients for success here.
Jane: So, in simple terms, the main implication is that we can't treat these foundation models as magic black boxes for energy pricing; we need to understand and incorporate the domain knowledge directly into their training or evaluation process to ensure they are reliable under real-world conditions.
Lu: It suggests a path forward where future work could involve developing better ways to measure and quantify exactly how much covariate information reduces forecast loss, which would help us design truly adaptive systems.
Meng: That’s a solid direction for future research; we need tools that tell us precisely where the model is falling short in terms of structural understanding, not just where the numerical error is high.
Lalam: I think this paper opens up exciting avenues for developing more nuanced AI agents that can dynamically adapt their reliance on external data based on the immediate market environment they are operating in.
Tom: It’s been fascinating watching how these authors dissect these complex dependencies, showing us exactly where the current state of TSFMs needs to mature before they become truly indispensable tools in this industry.
cs.LG, cs.SY, eess.SY
Submitted: 2026-07-02
Updated: 2026-10-07
Journal ref: ICML 2026 Foundation Models for Structured Data Workshop, 43rd International Conference on Machine Learning (ICML), Seoul, South Korea
Code: https://github.com/Nixtla/statsforecast
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 80/100
The gist: Time series foundation models (TSFMs) have shown strong zero-shot forecasting performance, but their generalization in covariate-driven, nonstationary settings is underexplored.
Key concepts
- Time Series Foundation Models (TSFMs)
- These are large models trained on vast amounts of time series data, capable of zero-shot forecasting. This study tests their ability to predict future electricity prices without needing specific training for that exact task, focusing on how well they handle real-world conditions.
- Covariate Support
- This refers to whether the model benefits from external information (covariates) like weather or market data. The research found that models incorporating relevant external data perform significantly better than those ignoring it, proving that domain knowledge and context are crucial for accurate forecasting.
- Probabilistic Performance (aQL)
- This measures how well the model predicts not just a single price point, but the entire range of possible outcomes. Average Quantile Loss (aQL) is used to assess this; better performance means the model provides more reliable confidence intervals for future electricity prices.
- Contamination Risk Mitigation
- Since TSFMs are pre-trained on general data, they might overfit to that training data when applied to a specific task like electricity pricing. The study uses two distinct datasets and careful selection criteria to ensure the evaluation reflects true generalization rather than just memorization of pre-training data.
Terminology
Summary
Time series foundation models (TSFMs) have shown strong zero-shot forecasting performance, but their generalization in covariate-driven, nonstationary settings is underexplored. This study proposes a two-dataset benchmarking framework for electricity price forecasting to mitigate contamination risk and enable fair evaluation of TSFMs across key aspects like point and probabilistic performance, tail behavior, and comparisons against domain-specific methods.
Benchmarking Framework
The study formalizes the electricity price forecasting task under a day-ahead market setup where the objective is to estimate the predictive distribution: P (Yt+δt:t+δt+H−1Ft). The researchers evaluate four TSFMs—Chronos 2, TimesFM 2.5, TabPFN-TS, and TOTO 1.0—examining both covariate-free and covariate-supported variants to assess the importance of exogenous information. Model selection was guided by three criteria: (i) covariates support; (ii) probabilistic forecast output; and (iii) accessibility and reproducibility, focusing on zero-shot rather than few-shot forecasting.
Model Comparison
The TSFMs are compared against three classes of baseline methods:
-
Classical time series models, including Seasonal Naïve, Auto-ARIMA, Auto-ETS, and Auto-Theta.
-
Task-trained deep learning approaches, such as DeepAR (Salinas et al., 2020), TFT (Lim et al., 2021), TiDE (Das et al., 2023a), TSMixer (Chen et al., 2023), DLinear, and NLinear.
-
Three domain-specific EPF methods: LEAR, DNN, and CING-LEAR.
Two-Dataset Evaluation Protocol
To mitigate contamination risk from pretraining corpora, the evaluation utilizes two complementary datasets: (i) GEFCom2014-P (a standardized benchmark), and (ii) GridStatus2025 (a newly curated dataset selected to minimize overlap with TSFM pretraining corpora). This design allows for a reasonably reliable assessment of model generalization while mitigating potential contamination effects,
as the real-world pretraining data for some models predates the evaluation period of GridStatus2025.
Performance Analysis and Key Findings
The results indicate that TSFMs are highly competitive and often outperform general-purpose baselines,
but their performance depends critically on covariate support.
Specifically, on the GridStatus2025 dataset, TSFMs generally outperform statistical and deep learning baselines, with covariate-supported variants (e.g., Chronos 2 w. and TabPFN-TS w.) consistently ranking among the strongest TSFMs across all metrics. Furthermore, a simple ensemble, constructed by averaging the forecasts of the best-performing TSFM (Chronos 2 w.) and domain-specific model (CING-LEAR) achieves the strongest overall performance,
suggesting that pre-trained TSFMs and domain-aware forecasting methods capture diverse and complementary predictive information.
Probabilistic and Tail Performance
Probabilistic performance is assessed using Average Quantile Loss (aQL). TSFMs achieve strong quantile-based performance compared with statistical and deep learning baselines,
with Chronos 2 w. and TabPFN-TS w. achieving the best aQL among TSFMs, again highlighting the importance of covariate support. Analysis of tail behavior reveals that median performance alone does not characterize forecasting behavior; several models exhibit substantial degradation in tail regions (especially lower tails)
compared to TSFMs and domain-specific methods. Statistical models struggle with lower-tail losses, whereas TSFMs produce more balanced tail forecasts across both extremes,
indicating improved distributional robustness. Finally, when evaluating robustness under price spikes (defined as observations falling below the 5th percentile or above the 95th percentile), CING-LEAR and Chronos 2 w. appear to better track price dynamics during most price spike events.
Covariate Dependence Assessment
The importance of exogenous information is statistically confirmed through Diebold-Mariano (DM) tests comparing covariate-incorporated variants against covariate-free variants. The results show that for Chronos 2, TimesFM 2.5, and TabPFN-TS, the one-sided alternative is supported by a rejection of the null hypothesis, indicating that covariate usage reduces forecast loss.
This reinforces the finding that accurate and robust performance in EPF is strongly tied to domain structure and task-specific inductive biases.
Conclusion
The study concludes that while TSFMs show competitive zero-shot and probabilistic forecast performance, "accurate and robust performance in EPF is strongly tied to domain structure and task-specific inductive biases.
Improvements for AI systems
As a fastidious and diligent researcher, I have thoroughly analyzed this paper, Evaluating Time Series Foundation Models for Electricity Price Forecasting: Contamination Risk, Distributional Shifts, and Covariate Dependence.
The core findings suggest that while Time Series Foundation Models (TSFMs) show promise in general forecasting, their success in complex electricity price forecasting (EPF) is critically dependent on domain-specific knowledge and the inclusion of relevant exogenous covariates.
Here are specific improvements to AI systems based on this research, detailing what the improved system can achieve:
- Improve Model Robustness via Covariate Integration
Based on Table 7 (Appendix A.5) and Section 3.1, TSFMs perform significantly better when they utilize relevant exogenous covariates (e.g., load, solar, gas prices). The Diebold-Mariano tests strongly reject the null hypothesis that covariate usage does not improve performance across MAE, RMSE, and aQL for models like Chronos 2 and TabPFN-TS.
Improved AI System Capability:
The system will transition from a generalist
forecasting model to an environmentally aware
one. It can now achieve superior accuracy in EPF by explicitly incorporating real-time or forecasted external data (weather, fuel prices, market features) rather than relying solely on historical price sequences. This allows the system to capture the structural and contextual dependencies of the energy market that are otherwise missing from purely time-series-based models.
- Enhance Distributional Accuracy and Risk Assessment
Section 3.2 highlights that median performance (Q(0.50)) is insufficient, especially in tail regions (e.g., Q(0.01) to Q(0.99)). Statistical methods struggle with low-price events, while TSFMs produce more balanced forecasts across both tails, suggesting improved distributional robustness.
Improved AI System Capability:
The system will gain risk-aware
forecasting capabilities. Instead of just predicting a single price point or a mean distribution, the system can output calibrated probabilistic forecasts (using quantile loss metrics like aQL) that accurately reflect the financial risk associated with extreme events—both high and low prices. This is crucial for grid operators and energy traders who need to understand the probability of encountering rare, volatile market conditions.
- Implement Contamination-Aware Deployment Strategy
The two-dataset benchmarking framework (GEFCom2014-P vs. GridStatus2025) strongly suggests that model performance is highly sensitive to the data distribution used for pretraining and evaluation, necessitating contamination-conscious evaluation protocols.
Improved AI System Capability:
The system deployment pipeline will be fortified with a rigorous validation layer that explicitly checks for data leakage
or temporal contamination from the pretraining corpus. The system can be deployed in production only if its performance metrics are validated across datasets that are temporally distinct from its training data, ensuring that the generalization capability is truly assessed rather than inflated by overlapping historical patterns.
- Develop Hybrid Ensemble Architectures for State-of-the-Art Performance
The results show that simple ensembles combining the best TSFM (e.g., Chronos 2 w.) and a domain-specific model (e.g., CING-LEAR) achieve the strongest overall performance, indicating complementary predictive information.
Improved AI System Capability:
The system will be architecturally designed as a Cognitive Ensemble.
It will utilize a meta-learning or gating mechanism to dynamically select and weigh outputs from specialized sub-models: one component for broad temporal pattern recognition (TSFM) and another for deep domain expertise (Domain-Aware Model). This hybrid approach allows the system to leverage the strengths of both foundation models and task-specific knowledge, leading to a performance ceiling that surpasses either component individually.
- Develop Spike/Nonstationarity Resilience Mechanisms
Section 3.3 demonstrates that models like Chronos 2 w. and CING-LEAR maintain better tracking during localized distributional shifts (price spikes) compared to other deep learning models, suggesting covariate support helps maintain stability under nonstationary conditions.
Improved AI System Capability:
The system will be engineered with an explicit Anomaly Response Module.
When the input data indicates a regime shift or potential price spike (based on statistical thresholds derived from the benchmark), the system can automatically switch its forecasting strategy. This could involve activating a more robust, domain-specific sub-module or dynamically adjusting its attention mechanisms to prioritize features that are known to be stable under stress, thereby minimizing forecast degradation during critical market volatility.
Abstract
Time series foundation models (TSFMs) have shown strong zero-shot forecasting performance, but their generalization in covariate-driven, non-stationary settings is underexplored. Electricity price forecasting (EPF) presents a challenging testbed due to complex temporal dependencies, distributional shifts, and strong reliance on structural and contextual information. We propose a two-dataset-benchmarking framework for EPF to mitigate contamination risk and enable fair evaluation of TSFMs. We examine key aspects of EPF including point and probabilistic forecasting performance, tail behavior, price spikes, and comparisons against domain-specific methods. We find that TSFMs are highly competitive and often outperform general-purpose baselines. Yet, their performance depends critically on covariate support, and they do not consistently surpass domain-specific methods tailored to EPF. Interestingly, simple ensembles of TSFMs and domain-specific methods appear to have significant potential, suggesting that the two approaches capture complementary predictive information.
Sources
- GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation
- Chronos-2: From Univariate to Universal Forecasting
- TSMixer: An All-MLP Architecture for Time Series Forecasting
- Toto: Time Series Optimized Transformer for Observability
- This Time is Different: An Observability Perspective on Time Series Foundation Models
- Long-term Forecasting with TiDE: Time-series Dense Encoder
- A decoder-only foundation model for time-series forecasting
- TimeGPT-1
- TempusBench: An Evaluation Framework for Time-Series Forecasting
- From Tables to Time: Extending TabPFN-v2 to Time Series Forecasting
- TS-Arena -- A Live Forecast Pre-Registration Platform
- Rethinking Evaluation in the Era of Time Series Foundation Models: (Un)known Information Leakage Challenges
- It's TIME: Towards the Next Generation of Time Series Forecasting Benchmarks
- fev-bench: A Realistic Benchmark for Time Series Forecasting
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks