Calibration Bets on the Past: Post-Training Quantization for Financial Time-Series Forecasting
Junyi Ye, Ivy Gateri Wanjiku
Montclair State University
cs.LG, q-fin.ST
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper presents a systematic study of activation calibration for post-training quantization (PTQ) in financial time-series forecasting, specifically for cross-sectional volatility forecasting on
Terminology
Summary
This paper presents a systematic study of activation calibration for post-training quantization (PTQ) in financial time-series forecasting, specifically for cross-sectional volatility forecasting on the S&P 500. The study covers seven representative neural architectures (DLinear, TSMixer, TimeMixer, vanilla Transformer, PatchTST, iTransformer, and SegRNN), eight walk-forward test years (2018–2025), and 560 independently trained models.
The key findings are:
-
8-bit quantization and weight-only 4-bit quantization preserve nearly all predictive skill. Under static W8A8, six of seven architectures lose no more than 0.0007 mean daily IC, corresponding to less than 0.2% of their FP32 signal. Weight-only W4 is also robust, with no architecture losing more than 2.5% of its FP32 IC.
-
Static 4-bit activation quantization causes substantial predictive losses. Under default abs-max calibration, PatchTST loses 11% of its FP32 signal, iTransformer loses 12%, TSMixer loses 20%, Transformer loses 25%, and both TimeMixer and SegRNN lose approximately 60%. The practical impact is significant: Transformer falls from an FP32 IC of 0.490 to 0.367 under default W4A4, placing it below the HAR baseline, while SegRNN falls from 0.458 to 0.187, below the persistence benchmark.
-
Much of the loss is recoverable through improved range selection. Replacing abs-max with percentile calibration recovers 53–94% of the degradation in the four most affected architectures. The range-recoverable share is 94% for Transformer, 80% for TSMixer, 73% for SegRNN, and 53% for TimeMixer. However, recurrent and multi-scale architectures retain substantial residual sensitivity.
-
The preferred activation range varies across market periods. The calibration–test mismatch produces two opposite failure modes: when the test year is more volatile than its calibration year (e.g., 2020 and 2022), narrow percentiles cost more than usual because they clip tail activations the calibration period underestimated; when the test year is calmer than its calibration year (e.g., 2021), abs-max becomes the costliest setting since it preserves resolution for extremes that no longer occur.
The paper also finds that under p99 calibration, damage outside the calibration envelope is 1.5–3.6× larger than inside for all four affected architectures, meaning the narrow range that helps on typical days clips more heavily once conditions turn extreme. A matched recalibration experiment shows that calibration–test mismatch accounts for roughly a quarter to half of the 2020 excess damage under p99, with the remaining loss persisting even when calibration matches the test year exactly.
The paper provides deployment guidance: measure default W4A4 damage before deployment; if damage is large, sweep percentile ranges on available data and compute range-recoverable and residual damage; if a single layer dominates the damage (as with Transformer's convolutional token embedding, which reproduces 72% of full-model W4A4 damage), keep that layer at 8 bits; if substantial residual damage remains, fall back to 8-bit activations, weight-only W4, or dynamic INT8. The guidance works without foresight of the deployment period, with the validation-selected percentile having mean regret of 0.007 IC and median regret of 0.002 IC relative to the oracle.
The central conclusion is that activation calibration is not a minor implementation detail but a first-order determinant of 4-bit deployment performance in financial forecasting.
Improvements for AI systems
Improvements to AI Systems:
-
Adaptive Activation Calibration for Time-Series Models: Implement percentile-based activation range selection (e.g., p99) instead of abs-max in PTQ pipelines for recurrent and multi-scale architectures (e.g., SegRNN, TimeMixer). This recovers 53–94% of predictive degradation, enabling reliable 4-bit deployment without retraining.
-
Volatility-Aware Range Selection: Add a calibration module that monitors market volatility (e.g., realized volatility of the target series) and dynamically adjusts the activation percentile—narrower percentiles during calm periods, wider during turbulent ones—to prevent tail clipping and resolution loss. This reduces calibration–test mismatch damage by up to 3.6×.
-
Layer-Specific Mixed-Precision Quantization: Automatically identify and retain 8-bit precision for outlier-sensitive layers (e.g., convolutional token embeddings in Transformers) that reproduce >70% of W4A4 damage. This preserves accuracy while keeping most of the model at 4-bit weights and activations.
-
Validation-Based Calibration Sweep with Regret Minimization: Replace oracle-based calibration with a validation-selected percentile sweep (mean regret 0.007 IC, median 0.002 IC). This allows deployment without foresight of future market conditions, making the system robust to unseen volatility regimes.
-
Failure-Mode Diagnostic for PTQ: Integrate a pre-deployment check that measures default W4A4 damage and decomposes it into range-recoverable vs. residual components. If residual damage is high, the system automatically falls back to weight-only W4 or dynamic INT8 activations, ensuring a minimum accuracy floor.
What the Improved AI System Can Do:
-
Deploy 4-bit quantized financial forecasting models on edge devices or low-latency trading systems with <2.5% loss in predictive skill (IC) for most architectures, and <12% loss even for the most sensitive ones, instead of 20–60% degradation.
-
Maintain forecasting accuracy across volatile market regimes (e.g., 2020, 2022) by adapting activation ranges in real time, avoiding catastrophic clipping of tail events.
-
Automatically optimize per-layer precision without manual tuning, reducing engineering effort while keeping Transformer-based models above baseline HAR performance.
-
Provide a reliable, foresight-free deployment pipeline that works on unseen future data, with near-oracle calibration performance (median regret of 0.002 IC).
-
Preserve ranking and risk-management decisions in cross-sectional volatility forecasting, enabling safer automated trading or portfolio optimization under memory and compute constraints.
Abstract
Financial forecasting models are typically developed in full precision, yet production deployment often requires low-precision inference to reduce memory and computational cost. Post-training quantization (PTQ) enables such deployment without retraining. However, reliable activation quantization requires calibration: activation ranges are estimated from historical data before deployment and then remain fixed during future inference. The importance of this deployment choice for financial forecasting remains poorly understood. We present a systematic study of activation calibration for PTQ in cross-sectional volatility forecasting on the S&P 500. Our evaluation covers seven representative neural architectures, eight walk-forward test years (2018-2025), and 560 trained models. We find that activation calibration has little effect at 8 bits but becomes the primary determinant of predictive performance at 4 bits. Under default absolute-maximum (abs-max) calibration, static 4-bit quantization of both weights and activations removes 11-62% of the full-precision mean information coefficient in affected architectures. Replacing abs-max with percentile calibration recovers 53-94% of this degradation in the four most affected architectures. The preferred activation range also varies across market periods. Narrow ranges improve resolution under typical market conditions but lose part of their advantage when test-period market dispersion exceeds the calibration history. These findings show that activation calibration is a first-class deployment decision for reliable 4-bit PTQ in financial forecasting. When substantial degradation remains, 8-bit activations or weight-only 4-bit quantization provide more robust deployment choices.
Sources
- PACT: Parameterized Clipping Activation for Quantized Neural Networks
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- A White Paper on Neural Network Quantization
- Quantizing Time-Series Models As Dynamical Systems: Trajectory-Based Quantization Sensitivity Score
- Assessing the Operational Viability of Foundation Models for Time Series Forecasting
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks