Market-Information-Aware Gated-LoRA of Foundation Models for Transferable Day-Ahead Electricity Price Forecasting
Hang Fan, Wei Wei, Shengwei Mei
cs.LG
Submitted: 2026-08-11
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: This paper proposes a market-information-aware adaptation framework that transfers the Chronos-2 time-series foundation model to day-ahead electricity price forecasting.
Terminology
Summary
This paper proposes a market-information-aware adaptation framework that transfers the Chronos-2 time-series foundation model to day-ahead electricity price forecasting. The authors identify a key gap: "Day-ahead electricity prices are formed by mapping anticipated supply–demand conditions into clearing prices. Therefore, future load, renewable output, reserve conditions, maintenance capacity, generator availability, and intertie schedules available before clearing are not auxiliary variables in a generic sense; they are part of the economic information set from which prices are formed." Existing supervised methods depend largely on market-specific historical data, limiting their use in newly established or data-scarce markets.
The framework has three main contributions. First, it constructs a multi-source market information (MSMI) interface that provides the pretrained backbone with a 7-day price context, historical covariates, known future information, and day-ahead market-clearing information over the forecasting horizon.
The interface includes "direct load, outgoing load, wind power, photovoltaic power, hydro power, nuclear power, upward and downward reserve-related quantities, ancillary-service quantities, maintenance capacity, generation capacity, and tie-line schedule." Second, it trains a source-domain gated low-rank adapter (LoRA) that updates approximately 1% of model parameters (about 1.21M of 120.7M total) without target-market labels. The gate scales the frozen source adapter according to market-state signals including reserve tightness, net load, renewable share, capacity tightness, and recent volatility.
Third, it adopts a leave-one-market-out (LOMO) protocol for evaluating cross-market transferability, where one market is held out as the target market
and no target-market training data are used in either zero-shot or source-market adaptation.
The experiments use four Chinese provincial day-ahead spot markets (Guangdong, Liaoning, Shandong, and Shanxi) with 15-minute resolution data, 7-day history context (L=672), and 1-day forecast horizon (H=96). Each held-out market is evaluated with 83 rolling daily windows. The results show that the proposed framework reduces the average MAE/RMSE by 6.24%/7.99% relative to market-information-aware zero-shot Chronos-2 and by 3.05%/3.52% relative to vanilla Source-LoRA.
Specifically, the MSMI interface reduces zero-shot average MAE from 86.25 (Core interface) to 79.60, Source-LoRA further reduces it to 76.98, and Gated LoRA achieves the best average MAE/RMSE of 74.63/137.87.
The paper reports several key findings. First, the largest gain comes from giving Chronos-2 access to market information available before day-ahead clearing,
with the MSMI interface providing a 7.7% relative improvement over the Core interface. Second, Source-domain LoRA provides a lightweight substitute for expensive domain re-pretraining.
Third, Market-state-gated LoRA makes the source adapter responsive to operating conditions and achieves the best average MAE/RMSE.
The control experiments show that a learned global scalar is not enough, and zero/random gate initializations remain close to the vanilla adapter,
while the reserve-initialized gate provides a useful economic inductive bias.
For probabilistic forecasting, the paper reports that Source-LoRA improves CRPS from 65.22 to 62.66, and gated LoRA further reduces it to 61.23.
However, the authors note a limitation: gated LoRA improves MAE and CRPS, but lowers PICP@80 and increases coverage deviation, so interval calibration remains an open issue.
The pooled PICP@80 decreases from 78.10% to 73.24% and then to 70.23%, with the macro-average absolute coverage deviation increasing from 5.00 to 6.76 to 9.77 percentage points.
The paper also examines a progressive adaptation spectrum. Source-domain LoRA reduces average MAE from 79.60 to 76.98 without target-market data. Direct few-shot tuning from pretrained Chronos-2 improves from 80.19 (5 target days) to 78.90 (30 target days), while using the source-domain adapter as a warm start gives a consistently stronger few-shot path
with five-seed averages decreasing from 77.11 (5 days) to 76.66 (30 days).
Statistical significance is evaluated using one-sided Diebold–Mariano tests on daily-window loss differentials. The pooled DM statistics show the MSMI interface significantly improves over Core (p < 0.01), Source-LoRA significantly improves over MSMI zero-shot (p < 0.05), and gated LoRA improves over Source-LoRA with marginal significance (p = 0.074). The paper notes that the gated-vs-Source statistical gain is marginal rather than decisive
for most markets, with the largest improvement appearing on Liaoning.
The paper acknowledges several limitations: future market-information variables are treated as available over the prediction horizon; the feature-interface ablation does not establish causal effects due to correlations among supply, demand, reserve, and generation-availability variables; gated LoRA lowers interval calibration; and the four-market corpus is modest relative to large-scale foundation-model pretraining datasets. Future directions include market-information robustness evaluation, market-adaptive architectures with dynamic rank allocation, and spike-aware training objectives.
The authors conclude with a staged practical workflow: market-information-aware zero-shot inference, source-domain LoRA adaptation, market-state calibration, and optional later target-market refinement.
They emphasize that for electricity price forecasting, foundation-model benchmarks should evaluate not only the backbone but also the information interface and adapter-control mechanism.
Improvements for AI systems
Improvements to AI systems:
-
Add a market-information interface layer to time-series foundation models. The AI system can ingest pre-clearing market data—load forecasts, renewable output, reserve requirements, maintenance schedules, generator availability, and tie-line plans—as structured inputs, not just historical price series. This enables the model to reason about supply–demand balance and price formation rather than treating price as an isolated sequence.
-
Implement source-domain LoRA adapters that update only 1% of parameters (e.g., 1.21M of 120.7M) using data from multiple related markets. The improved system can transfer to a new market with zero target-market labels, avoiding expensive full fine-tuning or re-pretraining, while preserving the backbone’s general temporal reasoning.
-
Add a market-state gate to the adapter, conditioned on interpretable signals like reserve tightness, net load, renewable share, capacity tightness, and recent volatility. The AI system can dynamically scale the frozen source adapter’s contribution based on current operating conditions, improving forecast accuracy when market stress or composition changes (e.g., high renewable penetration or tight reserves).
-
Adopt a leave-one-market-out evaluation protocol for cross-market generalization. The improved system can be benchmarked rigorously on held-out markets with no target training data, providing reliable estimates of zero-shot and adapted performance before deployment in new regions.
-
Enable a progressive adaptation workflow: start with market-information-aware zero-shot inference, then apply source-domain LoRA, then calibrate with market-state gating, and optionally refine with a few target-market days (e.g., 5–30 days) as a warm start. The system can deliver strong performance in data-scarce settings and improve further as small amounts of local data arrive.
-
Use reserve-initialized gate parameters as an economic inductive bias. The improved system can avoid poor local optima from zero/random initialization and achieve better average MAE/RMSE (e.g., 74.63/137.87) compared to vanilla adapters, with the gate responding meaningfully to market tightness.
-
Integrate probabilistic forecasting with CRPS optimization in the adapter training. The improved system can output calibrated prediction intervals, though interval calibration (PICP@80) remains a known weakness; the system can flag when coverage deviation is high and suggest fallback to simpler adapters for interval-based decisions.
-
Provide a multi-market information interface with 7-day context and 1-day horizon (e.g., L=672, H=96 at 15-minute resolution). The improved system can forecast day-ahead prices with 96 steps per day, using both historical covariates and known future schedules, achieving 6.24% lower MAE and 7.99% lower RMSE than zero-shot baselines.
What the improved AI system can do:
-
Forecast day-ahead electricity prices in new or data-scarce markets with no local training data, using only source-market adapters and pre-clearing market information.
-
Adapt to changing market conditions (e.g., reserve tightness, renewable share) by dynamically adjusting the adapter’s influence via the gate.
-
Provide both point forecasts and probabilistic forecasts (CRPS improved from 65.22 to 61.23) with a clear warning when interval calibration is unreliable.
-
Offer a staged deployment path: zero-shot first, then source adaptation, then market-state calibration, then optional few-shot refinement—each stage improving accuracy without requiring large target datasets.
-
Evaluate transferability across markets systematically, enabling safe deployment decisions before live use.
Abstract
Electricity price forecasting is crucial for market participants but remains difficult because prices are volatile, market-specific, and closely tied to anticipated system conditions. Existing supervised methods depend largely on market-specific historical data, limiting their use in newly established or data-scarce markets. This paper proposes a market-information-aware adaptation framework that transfers the Chronos-2 time-series foundation model to day-ahead electricity price forecasting. It first constructs a multi-source market information (MSMI) interface aligning 7-day price context with pre-clearing supply--demand, reserve, maintenance, generator-capacity, and intertie variables, and then trains a source-domain gated low-rank adapter (LoRA), updating about 1% of model parameters without target-market labels. The gate scales the frozen source adapter according to reserve-tightness and operating-state signals. A leave-one-market-out protocol is adopted for evaluating cross-market transferability. Experiments on four Chinese provincial day-ahead spot markets show that the proposed framework reduces the average MAE/RMSE by 6.24%/7.99% relative to market-information-aware zero-shot Chronos-2 and by 3.05%/3.52% relative to vanilla Source-LoRA. Experiments show that the gain is not reproduced by a learned global scalar or by random gate initialization, while the additional improvement over Source-LoRA is limited. These results suggest that market-structured inputs and state-dependent gated LoRA can provide a practical transfer path for data-scarce electricity markets.
Sources
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series Forecasting
- Chronos-2: From Univariate to Universal Forecasting
- Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting
- PriceFM: Foundation Model for Probabilistic Electricity Price Forecasting
- A Time Series is Worth 64 Words: Long-term Forecasting with Transformers
- iTransformer: Inverted Transformers Are Effective for Time Series Forecasting
- Are Transformers Effective for Time Series Forecasting?
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks