Quantile Adaptive Temperature Scaling for Confidence Calibration
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Quantile Adaptive Temperature Scaling for Confidence Calibration".
Jane: Deep neural networks often produce poorly calibrated confidence estimates, overstating their certainty even when predictions are incorrect, which makes calibration vital for reliable risk estimation in high-stakes domains.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Okay, moving on to the title and who wrote this paper, "Quantile Adaptive Temperature Scaling for Confidence Calibration." It's clear that the focus here is on making confidence estimates better by adapting them based on quantiles.
Jane: The authors are Omprakash Chakraborty, Leo Fillioux, Ismail Ben Ayed, and Jose Dolz from various institutions including ÉTS Montréal and Université Paris-Saclay. They bring a good mix of theory and practical application to this calibration problem.
Lu: It's interesting seeing researchers from such diverse backgrounds collaborate on a post-hoc method that addresses the structural limitations of existing techniques like standard Temperature Scaling.
Meng: I wonder how many different model architectures these authors tested, because knowing if this works across different model types is crucial for me to assess its practical utility.
Lalam: For me, the fact that they are focusing on a simple and efficient post-hoc method that corrects the heterogeneity of miscalibration feels very promising for deployment in real-world systems.
The paper's summary: Tom: The paper summarizes the core idea behind Quantile Adaptive Temperature Scaling for Confidence Calibration, which is introducing a sample-wise temperature function that depends on the prediction's empirical confidence quantile.
Jane: In simpler terms, instead of applying one uniform adjustment to all the model outputs, QaTS lets the model adjust its temperature based on how confident it is for that specific prediction.
Lu: This approach breaks away from the assumption that a single smooth function can fix everything globally; by conditioning on quantiles, they are able to correct errors in different parts of the confidence spectrum more effectively.
Meng: So, if I understand correctly, the methodology involves mapping raw confidence values into a quantile space using an empirical cumulative distribution function before defining the temperature adjustment? That sounds like a solid mathematical step.
Lalam: Exactly, and that mapping allows them to introduce learnable parameters 'a' and 'b' into the temperature formula, which means it’s not just a fixed rule; it learns how to best adapt for each prediction.
The paper's improvements: Tom: The main improvement discussed in "Quantile Adaptive Temperature Scaling for Confidence Calibration" is moving away from a single global rescaling function and instead using a sample-wise temperature function.
Jane: By defining the temperature as a monotone linear function of (one - q(x)), where q(x) is the empirical confidence quantile, they are directly targeting where the miscalibration is most severe in different regions <ref:2606.21749#pg0>.
Lu: This rank-based formulation is key because it sidesteps the global smoothness assumption that plagues methods like standard Temperature Scaling, allowing corrections to vary systematically across the entire spectrum.
Meng: So, when they talk about correcting overconfident predictions at low quantiles while sharpening high quantiles, that means they are specifically addressing the known structural differences in where errors occur.
Lalam: And because this method inherits the rank-based stability of quantiles, it makes QaTS substantially less sensitive to global logit scaling effects that often happen when data distribution shifts occur.
Conclusion: Tom: So, to wrap up "Quantile Adaptive Temperature Scaling for Confidence Calibration," the paper shows how adapting temperature based on confidence quantiles directly targets the heterogeneity of miscalibration that standard methods miss.
Jane: It really suggests that this quantile reparameterization is a powerful way to align calibration with a reparameterized Expected Calibration Error objective, leading to better reliability in complex settings.
Lu: The implication here is that for models in areas like healthcare or autonomous driving, this method offers a way to get more accurate uncertainty quantification by explicitly modeling the confidence distribution structure.
Meng: From a practical standpoint, it means we can deploy post-hoc calibration that doesn't require retraining and is robust enough to handle covariate shifts without breaking its calibration effectiveness.
Lalam: I think the main impact is that we are getting much more reliable risk estimation for AI systems, because we're not just getting one average score; we’re getting a calibrated view of uncertainty tailored to the prediction's confidence level.
ÉTS Montréal, Canada · Université Paris-Saclay · CentraleSupélec · Gustave Roussy INSERM CDSU IHU PRISM
cs.CV
Submitted: 2026-06-19
Updated: 2026-10-02
Importance score: 88/100
The gist: Deep neural networks often produce poorly calibrated confidence estimates, overstating their certainty even when predictions are incorrect, which makes calibration vital for reliable risk estimation
Key concepts
- Quantile Reparameterization
- This technique maps raw prediction confidences into a uniform quantile axis using the empirical cumulative distribution function. It sorts all sample confidences and calculates where a specific sample's confidence falls within that sorted distribution, allowing the method to operate directly in quantile space rather than raw logit space.
- Sample-wise Temperature Function
- QaTS defines a temperature T(x) based on the prediction's empirical quantile q(x). The function is defined as T(x) = a * (1 - q(x)) + b, where 'a' and 'b' are learnable positive parameters. This allows the scaling factor to change systematically depending on how confident or uncertain a specific prediction is.
- Heterogeneous Miscalibration
- The paper observes that errors in neural networks are not uniform across all confidence levels; they vary significantly. QaTS addresses this by tailoring the temperature adjustment to these specific regions—softening overconfident predictions while sharpening those with high uncertainty—which standard methods fail to do effectively.
Terminology
Summary
Deep neural networks often produce poorly calibrated confidence estimates, overstating their certainty even when predictions are incorrect, which makes calibration vital for reliable risk estimation in high-stakes domains. Quantile-Adaptive Temperature Scaling (QaTS) is introduced as a simple and efficient post-hoc calibration method that adapts the temperature as a function of a prediction’s empirical confidence quantile, effectively correcting the heterogeneous structure of miscalibration that standard methods fail to address.
The gist
QaTS introduces a sample-wise temperature function, defined as a monotone linear function of (1 - q(x)), which conditions the temperature on the sample’s confidence quantile rather than its raw logit magnitude. This rank-based formulation breaks the global-smoothness assumption and enables corrections that vary systematically across the confidence spectrum, aligning with a reparameterized Expected Calibration Error (ECE) objective.
Problem Definition and Motivation
The core motivation stems from the observation that miscalibration is far from uniform across the whole confidence spectrum; specifically, in standard regimes, largest correctness–confidence gaps occur in the lowest confidence quantiles where uncertainty is highest, while in long-tailed regimes, dominant errors shift toward medium-confidence quantiles. Existing temperature scaling (TS) methods implicitly assume that miscalibration can be corrected by a single global rescaling function applied uniformly across the confidence spectrum. This structural rigidity means they cannot effectively adapt their corrections to regions where miscalibration is more severe, leading to degraded performance under challenging scenarios like distributional shifts.
Quantile Reparameterization and Temperature Function
QaTS addresses this by operating directly in quantile space. It first defines an empirical quantile function, mapping raw confidences into a uniform quantile axis using the empirical cumulative distribution function:
-
Sort the confidences of all calibration samples:
conf(x(1)) ≤ conf(x(2)) ≤ · · · ≤ conf(x(N)).
-
Define the empirical quantile function as:
F conf(u) = (1/N) sum over i of 1 if conf(x(i))≤ u.
-
The empirical quantile for a sample is then computed as:
q(x) = F conf(conf(x)).
This quantile-based parameterization allows the introduction of a sample-wise temperature function, defined as:
**"T(x) = a **
(1 - q(x)) + b, where a and b are two positive learnable parameters."
Learning the Temperature Parameters
The parameters 'a' and 'b' are optimized using smooth surrogates like the Negative Log Likelihood (NLL), which is a strictly proper scoring rule. The objective is to minimize the NLL over calibration samples:
"min a, b > 0 - (1/N) sum over i of 1 sum over k of K y(i) k log (tilde p k(x(i);a,b))."
This optimization yields parameters that adaptively adjust the temperature based on each prediction’s confidence quantile, directly targeting the heterogeneity revealed by the quantile reparameterization and thereby improving ECE.
The resulting scaled probabilities are then used for inference:
tilde p k(x) = exp (l k(x)/T(x)) / sum over j of 1 exp (l j(x)/T(x)).
Benefits and Robustness
The primary benefit is that QaTS explicitly aligned with minimizing Eq. (7)
by modifying the confidence precisely in the regions where Eq. (7) shows the largest contribution to ECE, specifically by softening low-quantile (overconfident) predictions while sharpens high-quantile ones. Furthermore, mapping confidences through the calibration CDF yields a monotone transform, making QaTS stable under monotone transformations of the logits
and substantially less sensitive to global logit scaling effects that often accompany covariate shift.
This makes it robust across domains and scenarios.
Experimental Validation
QaTS was validated across 162 different settings, including standard computer vision (CIFAR-10, ImageNet), long-tailed distributions (CIFAR-100-LT), distributional shifts (CIFAR-100-C), medical imaging tasks (HAM10000, MedMNIST2D), and text classification. Across all these settings, QaTS consistently reduces calibration bias and error
and yields superior uncertainty calibration than state-of-the-art post-hoc calibration methods,
often achieving the lowest Expected Calibration Error (ECE) or Adaptive Expected Calibration Error (AECE). The results confirm that QaTS preserves accuracy while improving reliability.
Compatibility with Training-Time Calibration
QaTS also serves as a complementary post-hoc approach to training objectives.
Improvements for AI systems
Based on the scientific paper, here are the specific improvements for AI systems that could be implemented:
) The proposed Quantile-Adaptive Temperature Scaling (QaTS) method should be integrated into any post-hoc calibration pipeline of deep neural networks. This involves modifying the inference phase to use a sample-wise temperature function, which is a monotone linear function of the empirical confidence quantile, rather than applying a single global temperature scaling factor to all logits.
] The improved system can provide more reliable and trustworthy confidence estimates for predictions across diverse scenarios (standard classification, long-tailed distributions, and distributional shifts) without requiring any retraining or modification of the underlying model weights.
[1] Specifically, QaTS enables the model to:
-
Identify regions of overconfidence (low-quantile predictions) and apply a higher temperature to these logits. This effectively softens the prediction, reducing its confidence score and aligning it with empirical correctness frequencies.
-
Preserve high-confidence predictions (high-quantile predictions) by applying minimal or lower temperature adjustments, thereby avoiding unnecessary distortion of reliable estimates.
-
Achieve superior Expected Calibration Error (ECE) reduction compared to state-of-the-art methods like Temperature Scaling (TS), Isotonic Regression (IR), and AdaTS, especially in challenging scenarios such as class imbalance and distributional shifts.
[2] The improved system can perform more robust risk estimation in high-stakes domains where miscalibration is costly, such as autonomous driving or critical infrastructure. Because QaTS is robust to covariate shifts (as it relies on rank-based information), it maintains its calibration effectiveness even when test samples differ significantly from the calibration set, unlike methods that rely on absolute confidence values.
[3] The improved system can be deployed with simpler inference overhead compared to complex feature-space or input-dependent temperature scaling methods (like AdaTS) because it only requires a single learnable function defined over the quantile space, making it efficient for real-time applications.
[4] The improved system can maintain high discriminative performance while significantly improving calibration. The paper demonstrates that QaTS outperforms non-accuracy preserving methods like FeatClip in terms of accuracy across most settings, ensuring that the model's ability to correctly identify classes is not compromised by the calibration technique.
[5] The improved system can be seamlessly integrated with existing training-time regularization objectives (like Focal Loss or Brier Loss) without requiring modifications to the original optimization process, acting as a powerful complementary post-hoc tool to enhance overall model reliability.
[6] The improved system can provide more accurate uncertainty quantification in complex, real-world environments. By aligning predicted confidence with actual correctness across the entire distribution of confidences (as measured by AECE), QaTS offers a superior measure of predictive uncertainty compared to traditional ECE metrics, which is crucial for tasks requiring precise risk assessment.
Sources
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models