Retrieval-Corrected Conformal Prediction for Time Series

arXiv:2608.10553 · cs.LG, cs.AI · Submitted 2026-08-11 · Read on arXiv

Sangjin Jin, Kangmin Kim, Junhyeong Lee, Yongjae Lee

Ulsan National Institute of Science and Technology · LinqAlpha

cs.LG, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: Accepted at the 35th ACM International Conference on Information and Knowledge Management (CIKM '26), Rome, Italy

Code: https://github.com/jinsaaang/rccp

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: Retrieval–Corrected Conformal Prediction (RCCP) is a retrieval-augmented calibration method for time series prediction intervals.

Terminology

Summary

Retrieval–Corrected Conformal Prediction (RCCP) is a retrieval-augmented calibration method for time series prediction intervals. The paper states: "RCCP builds an asymmetric interval from retrieved one-sided residuals and calibrates its normalized retrieval error with a scalar conformal correction. Thus, retrieval provides local residual evidence, while conformal correction determines the final scale needed for coverage."

The method addresses a limitation of existing time series conformal prediction methods: "Recent time series CP methods improve local calibration using recent, weighted, or localized residuals. Yet local calibration can remain indirect, since broad residual weighting or additional adaptation procedures may dilute the evidence most relevant to the current prediction. RCCP instead selects similar past residuals as local evidence and then corrects the coverage error left by retrieval."

The approach works in two stages. First, RCCP retrieves similar past prediction contexts and their realized residuals from a time ordered knowledge base, then builds an asymmetric local interval from positive and negative residuals. Second, "It instead evaluates the remaining retrieval error on calibration data by comparing the realized residual with the retrieved upper or lower scale. The resulting normalized retrieval error is calibrated into a scalar correction factor for test time intervals."

The paper's main contributions are: applying residual retrieval to conformal prediction as a direct alternative to broad weighting over calibration samples and additional training procedures; proposing RCCP as a framework that combines asymmetric retrieved intervals with scalar conformal correction; providing a coverage gap bound and an asymptotic coverage result for the RCCP interval, based on stability of the normalized retrieval error distribution; and showing empirically that RCCP improves the coverage and interval efficiency tradeoff with fewer large misses while keeping calibration cost low.

The theoretical analysis establishes that "the bound separates calibration-test score mismatch from empirical correction error. Retrieval affects the interval through the sharpness and stability of Cet, while coverage is controlled through the normalized retrieval error Bt. The asymptotic result shows that as the calibration-test score mismatch and empirical correction error vanish, Ct(bc) attains the target coverage asymptotically."

Empirically, across four benchmark datasets (Air, Solar, Wind, Electricity) and two backbone forecasters (LSTM and Transformer), RCCP avoids empirical undercoverage at alpha = 0.1 and achieves the best or tied-best Winkler scores. The paper reports that RCCP’s advantage does not come from producing the narrowest intervals, but from reducing the miscoverage penalty, and that RCCP widens intervals in high-error deciles and maintains stronger coverage where prediction errors are largest, with width adaptivity ratios reaching 2.30 on Air and 20.86 on Solar, exceeding all baselines.

Ablation studies show that conformal correction mainly controls coverage and asymmetric residual scaling mainly controls efficiency, with the full design justified because correction ensures reliable coverage, while asymmetry improves interval efficiency. The correction factor is robust across target miscoverage levels (alpha ∈ 0.05, 0.10, 0.15), and coverage remains stable across variants of retrieval key construction and distance metrics. The method is also robust to neighborhood size, with the default K = 64 provid[ing] a stable balance between local adaptivity and interval quality.

The paper concludes: These results suggest that retrieval and correction provide a simple and scalable route to locally informative and calibrated uncertainty estimates for time series forecasting. A stated limitation is its dependence on the quality of the retrieval representation, noting that interval efficiency may degrade when the embedding does not capture error-relevant similarity.

Improvements for AI systems

Improvements to AI Systems:

  1. Direct Local Evidence Retrieval for Uncertainty Calibration: Replace broad residual weighting or additional training-based adaptation in time series conformal prediction with a retrieval step that selects the most similar past prediction contexts and their realized residuals. This directly targets the local error distribution without diluting evidence, improving interval sharpness and reducing large misses.

  2. Asymmetric Interval Construction from Retrieved Residuals: Build prediction intervals using separate upper and lower scales derived from positive and negative retrieved residuals, rather than symmetric intervals. This captures skewness in forecast errors (e.g., asymmetric volatility in energy or demand data), improving efficiency without sacrificing coverage.

  3. Scalar Conformal Correction on Normalized Retrieval Error: After retrieval, compute a single scalar correction factor by calibrating the normalized retrieval error (comparing realized residuals to retrieved scales) on a validation set. This corrects any residual coverage gap left by retrieval, ensuring reliable coverage without retraining or complex multi-step adaptation.

  4. Stability-Aware Retrieval for Robust Coverage: Use the theoretical bound to guide retrieval key design (e.g., embedding choice, distance metric) by monitoring the stability of the normalized retrieval error distribution. This improves robustness to representation quality, preventing efficiency degradation when embeddings poorly capture error-relevant similarity.

  5. Adaptive Width Control in High-Error Regimes: Automatically widen intervals in deciles where prediction errors are large (as demonstrated by width adaptivity ratios up to 20.86), while keeping intervals tight in low-error regions. This reduces the miscoverage penalty and improves the coverage–efficiency tradeoff, particularly for volatile time series like solar or wind power.

  6. Low-Cost, Training-Free Calibration Pipeline: Implement RCCP as a post-hoc, retrieval-augmented calibration layer that requires no additional training of the forecaster. This enables rapid deployment on existing LSTM or Transformer forecasters, with calibration cost kept low (only retrieval and scalar correction), making it scalable to large-scale forecasting systems.

What the Improved AI System Can Do:

  • Produce prediction intervals for time series that achieve target coverage (e.g., 90%) without systematic undercoverage, even with non-stationary or heteroscedastic errors.

  • Dynamically adjust interval width based on local error evidence, providing tighter intervals in stable periods and wider intervals during high-uncertainty events (e.g., demand spikes, weather shifts).

  • Maintain high interval efficiency (lower Winkler scores) by reducing the frequency and magnitude of large misses, outperforming baselines on datasets like Air, Solar, Wind, and Electricity.

  • Operate as a drop-in module for any time series forecaster, requiring only a time-ordered knowledge base of past contexts and residuals, with no retraining or architectural changes.

  • Provide theoretical guarantees on coverage asymptotically, with a bound that separates calibration–test mismatch from empirical correction error, enabling reliable deployment in safety-critical forecasting (e.g., energy grid management, logistics).

Abstract

Conformal prediction (CP) provides distribution-free prediction intervals for fixed forecasters, but its standard calibration procedure is often inefficient for time series data, where forecast errors are temporally dependent and change across time and operating conditions. Recent time series CP methods improve local calibration using recent, weighted, or localized residuals. Yet local calibration can remain indirect, since broad residual weighting or additional adaptation procedures may dilute the evidence most relevant to the current prediction. This motivates a simple retrieval and correction strategy that selects similar past residuals as local evidence and then corrects the coverage error left by retrieval. In this paper, we propose Retrieval--Corrected Conformal Prediction (RCCP), a retrieval-augmented calibration method for time series prediction intervals. RCCP builds an asymmetric interval from retrieved one-sided residuals and calibrates its normalized retrieval error with a scalar conformal correction. Thus, retrieval provides local residual evidence, while conformal correction determines the final scale needed for coverage. We provide a coverage-gap bound based on the stability of the normalized retrieval error distribution. Across standard benchmarks and backbone forecasters, RCCP attains the target coverage in every setting and achieves the lowest Winkler scores, with fewer severe misses. RCCP also achieves low calibration and inference overhead, showing that retrieval-corrected calibration is an effective and scalable approach to uncertainty quantification in time series forecasting. Code is available at https://github.com/jinsaaang/rccp.

Sources

Related papers