Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration
cs.LG, cs.CV
Submitted: 2026-09-01
Updated: 2026-09-02
Code: https://github.com/labhai/CORD
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 prediction.
Terminology
Abstract
Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 prediction. Accuracy captures only the net effect of these changes on correctness, not how often predictions change; the Top-1 Prediction Change Rate (TPCR) instead measures this frequency. We propose Calibrator-Output Repair for Top-1 Decision Preservation (CORD), the first post-fit adapter to impose exact prediction preservation by repairing the full calibrated probability vector. From the original and calibrated outputs alone, CORD determines the mass assigned to the original top-1. The calibrated conditional distribution allocates the remaining mass over the other classes, yielding a repaired vector whose own argmax recovers the original prediction. On the calibration split, CORD coordinates the repaired masses to retain the calibrated outputs' mean mass on original predictions whenever attainable. The adapter alters neither the fitted calibrator nor its direct output, fits no additional supervised map, and requires no user- or validation-tuned hyperparameter. Across CIFAR-10/100 and ImageNet-1K, CORD attains zero TPCR by construction and lowers mean ECE, NLL, and Brier relative to the corresponding direct outputs in every dataset; paired gains persist under distribution shift and across calibration-set sizes. CORD thus removes the preservation constraint from calibrator fitting and assigns exact recovery of the original decision to subsequent output repair. Our code is available at https://github.com/labhai/CORD.
Sources
- Probability Bounding: Post-Hoc Calibration via Box-Constrained Softmax
- Structured Matrix Scaling for Multi-Class Calibration
- Quantile Adaptive Temperature Scaling for Confidence Calibration
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
- Rethinking Post-Hoc Calibration in Semantic Segmentation
- Wide Residual Networks
- Mix-n-Match: Ensemble and Compositional Methods for Uncertainty Calibration in Deep Learning
- Instance-Wise Monotonic Calibration by Constrained Transformation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks