Calibration of Ordinal Regression Networks
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Calibration of Ordinal Regression Networks".
Jane: The paper was written by Daehwan Kim, Haejun Chung and Ikbeom Jang from Hanyang University and Hankuk University of Foreign Studies.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone. Today we’re looking at a paper that’s got a mouthful of a title: “Calibration of Ordinal Regression Networks.” I’m here with Jane, and we’re both pretty excited about this one because it tackles something that doesn’t get enough attention.
Jane: Absolutely, Tom. So, when we talk about ordinal regression, we’re talking about tasks where the labels have a natural order. Think age groups, like “child, teen, adult, senior,” or disease severity, like “mild, moderate, severe.” It’s not just about getting the right answer; it’s about understanding that “moderate” is closer to “mild” than it is to “severe.”
Tom: Right, and the paper’s main point is that most models are really good at predicting the label but really bad at telling you how confident they are. That’s the calibration part. If a model says it’s ninety percent sure, it should be right ninety percent of the time. But in these ordered tasks, that often falls apart.
Jane: And that’s a huge deal. Imagine a medical AI looking at a scan and saying, “I’m ninety-five percent sure this is stage two.” If it’s actually only right seventy percent of the time at that confidence level, that’s dangerous. The authors point out that existing methods either focus on getting the order right or on calibration, but rarely both at the same time.
Tom: Exactly. They call this the gap. You have ordinal-focused losses that make the probabilities look nice and unimodal—you know, peaking at the right answer and trailing off smoothly—but they’re often overconfident or underconfident. Then you have calibration-focused losses that fix the confidence but ignore the order, so you get predictions that make no sense, like being super confident about “senior” and then second most confident about “child.”
Jane: So the title really captures the dual challenge. It’s not just about making a classifier; it’s about making a classifier that respects the order of the labels and tells you the truth about its uncertainty. The authors are basically saying, “Hey, we can do both, and here’s how.”
Tom: And that’s the hook for us. We’ll get into the actual method they propose in a bit, but the implication is clear. For high-stakes fields like medicine, autonomous driving, or even content rating systems, having a model that’s both accurate and honest about its confidence is a game-changer.
Jane: It really is. And it’s one of those problems that feels like it should have been solved already, but the paper shows that the standard tools just don’t cut it for ordered data. So, stick around. Next, we’re going to break down their actual solution and why it works so well.
Summary: Tom: So, Jane, we’ve set the stage. The paper “Calibration of Ordinal Regression Networks” is about fixing the confidence problem in ordered classification. But what’s their actual idea? Let’s get into the summary.
Jane: Right. The core issue they start with is that the standard cross-entropy loss, which is the workhorse for classification, treats every class as if it’s completely unrelated. For ordered labels, that’s a mistake. It pushes the model to be overly confident about the exact label and ignores the fact that “moderate” is closer to “mild” than “severe.”
Tom: So they propose a new loss function called ORCU. Catchy name. It stands for Ordinal Regression loss for Calibration and Unimodality. And it’s got two main parts.
Jane: Two parts, exactly. The first part uses something called soft ordinal encoding. Instead of a one-hot vector where only the true label is a one and everything else is a zero they create a soft target distribution. The true label gets the highest value, but neighboring labels get some probability too, based on how close they are. This way, the model learns that being off by one is better than being off by three.
Tom: That makes sense. It’s like telling the model, “Don’t just memorize the answer; understand the neighborhood.” But that alone isn’t enough, right? Because that could make the model too timid, too underconfident.
Jane: Exactly. That’s the second part. They add an ordinal-aware regularization term. This term actively checks the differences between the logits, which are the raw scores before the softmax. It enforces a rule: the logits should decrease as you move away from the true label. If they don’t, the loss gets bigger.
Tom: And this is where the clever bit comes in. The regularization isn’t a flat penalty. It’s adaptive. If the model is already doing a good job and the logits are nicely ordered, the penalty is small. But if the model is about to violate the order, the penalty ramps up sharply.
Jane: Right. It’s like a guardrail that only kicks in when you’re about to go off the road. This helps the model stay unimodal, meaning the probabilities peak at the true label and fall off smoothly, which is a natural assumption for ordered data.
Tom: And the results? They tested this on four datasets, including age estimation from photos and medical images for diabetic retinopathy and ulcerative colitis. The paper shows that ORCU consistently gets much lower calibration errors than both the ordinal-focused and the calibration-focused baselines.
Jane: And it does this without sacrificing accuracy. In fact, on some metrics like quadratic weighted kappa, which measures how well the model respects the order, ORCU actually does better than the baselines. So it’s not a trade-off; it’s a win-win.
Tom: That’s the dream scenario. Better confidence and better ordered predictions. So, the summary is that they’ve built a loss function that teaches the model to respect the order of the labels and be honest about its uncertainty, all in one go.
Jane: And that’s a big step forward. Next, we’re going to dig into the improvements this suggests over existing methods and why the details of that regularization term matter so much.
Improvements: Tom: We’re back with the paper “Calibration of Ordinal Regression Networks.” Jane, we’ve covered the big idea. Now let’s get into the nitty-gritty of what makes ORCU better than what came before.
Jane: Good place to start. The paper compares ORCU against a whole lineup of existing losses. On the ordinal side, you have things like SORD, which uses soft encoding, and CDW-CE, which adds a distance penalty. On the calibration side, you have label smoothing and focal loss variants.
Tom: And the key improvement is that none of those methods do both jobs well. For example, SORD gets the order right—it produces very high unimodality, almost one hundred percent—but its calibration errors are still high. It’s underconfident because the soft targets are too diffuse.
Jane: Right. And the calibration-focused methods, like label smoothing, can get the overall confidence down, but they don’t respect the order. The paper shows that they often produce non-unimodal distributions, which means the model is saying something like “I’m most confident about class two but my second most confident is class five.” That’s nonsense for ordered data.
Tom: So ORCU’s improvement is that it explicitly ties the calibration to the ordinal structure. The regularization term isn’t just about making the confidence match accuracy; it’s about making the confidence match accuracy *and* ensuring the probabilities fall off in the right direction.
Jane: And the details matter. They use a log-barrier function for the regularization. That sounds fancy, but the idea is simple. When the logit difference is far from the violation point, the penalty is gentle. As it gets closer to violating the order, the penalty grows logarithmically, which is a strong push to correct course.
Tom: It’s like a spring that gets stiffer the more you stretch it. That’s a smart way to avoid over-regularizing the easy samples while still catching the hard ones.
Jane: Exactly. And they also show that this works across different model sizes. They tested it with ResNet-thirty-four ResNet-fifty and ResNet-one hundred one and the improvements hold. That’s a good sign for generalizability.
Tom: Another improvement is in the feature embeddings. The paper includes t-SNE visualizations showing that ORCU produces much more ordered clusters in the feature space. The classes line up in a neat progression, whereas the baselines have more scattered, overlapping clusters.
Jane: That’s a visual confirmation that the model is learning the underlying structure, not just memorizing the labels. And that’s important for real-world use, where you might encounter new data that’s slightly different from the training set.
Tom: So the improvements are clear. Better calibration, better order awareness, and better feature representations. All from a single loss function.
Jane: And that’s the kind of practical improvement that could make a real difference in deployment. Next up, we’re going to wrap up with our final thoughts and what this means for the future of reliable AI.
Conclusion: Tom: Alright, we’re wrapping up our discussion on “Calibration of Ordinal Regression Networks.” Jane, what’s the big takeaway for our listeners?
Jane: The big takeaway is that we now have a way to make ordinal classifiers both accurate and trustworthy. The ORCU loss function is a single, elegant solution that addresses the two biggest problems in this space: respecting the order of the labels and providing reliable confidence estimates.
Tom: And that’s not just a nice-to-have. In fields like medicine, where a model’s confidence can influence a diagnosis, or in any system that triages based on severity, having calibrated confidence is critical. This paper gives us a tool to build those systems more safely.
Jane: I also appreciate that the paper is honest about its limitations. They focused on computer vision datasets, but the method is general enough that it could be applied to text or audio tasks. That’s a natural next step for future research.
Tom: And the future work they mention is about tackling the remaining miscalibration at extreme confidence levels. So there’s still room to improve, but this is a solid foundation.
Jane: For sure. It’s one of those papers that makes you wonder why we weren’t doing this all along. The idea is simple, but the execution is thorough, and the results speak for themselves.
Tom: So, we’re saying goodbye to this paper. It’s been a great discussion. Thanks to everyone listening, and we’ll be back soon with the next one.
Jane: Until then, keep questioning the confidence of your models. It might just save a life.
Daehwan Kim, Haejun Chung, Ikbeom Jang
Hanyang University · Hankuk University of Foreign Studies
cs.LG, cs.CV
Submitted: 2026-08-14
Updated: 2026-08-17
Code: https://github.com/Jonathan-Pearce/calibration_library
Project page: https://talhassner.github.io/home/projects/Adience/Adience-data.html
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 51/100
The gist: The paper "Calibration of Ordinal Regression Networks" by Daehwan Kim, Haejun Chung, and Ikbeom Jang addresses the dual challenges of calibration and ordinal relationship modeling in ordinal
Key concepts
- Ordinal Regression
- A type of task where labels have a natural order, such as disease severity (mild to severe) or age groups. It requires the model to understand that the distance between categories matters, not just getting the correct label.
- Calibration
- The ability of a model to accurately report its confidence. If a model states it is 90% sure of an outcome, it should be correct approximately 90% of the time. Poor calibration is dangerous in high-stakes applications.
- ORCU Loss Function
- A novel loss function proposed by the paper. It combines two parts: soft ordinal encoding (to respect label order) and an adaptive regularization term (to ensure accurate confidence estimates), solving both challenges simultaneously.
- Unimodality
- A property where a probability distribution peaks at the true answer and falls off smoothly as you move away from it. This is a natural assumption for ordered data, indicating the model's confidence decreases gracefully with distance.
Terminology
Summary
The paper Calibration of Ordinal Regression Networks
by Daehwan Kim, Haejun Chung, and Ikbeom Jang addresses the dual challenges of calibration and ordinal relationship modeling in ordinal regression tasks. The authors state: "Recent studies have shown that deep neural networks are not well-calibrated and often produce over-confident predictions. The miscalibration issue primarily stems from using cross-entropy in classifications, which aims to align predicted softmax probabilities with one-hot labels. In ordinal regression tasks, this problem is compounded by an additional challenge: the expectation that softmax probabilities should exhibit unimodal distribution is not met with cross-entropy."
The paper notes that the ordinal regression literature has focused on learning orders and overlooked calibration
and that existing approaches have neglected mainly the joint consideration of calibration and ordinal relationships, leaving a critical gap in addressing these interdependent aspects effectively.
The authors claim this is the first study to explicitly address both calibration and ordinality in ordinal regression.
The proposed solution is a novel loss function called ORCU (Ordinal Regression loss for Calibration and Unimodality), which jointly enforces calibration and ordinal relationships.
The authors explain: ORCU explicitly encodes the ordering of classes and incorporates an ordinal-aware regularization term to align predicted probabilities with the inherent class order while preserving reliable confidence estimates.
The method consists of two main components:
1. Ordinal Soft Encoding: The authors adopt SORD encoding (Diaz & Marathe, 2019), which replaces one-hot encoding with a soft encoding that distributes values across classes based on inter-class distances, thereby mitigating overconfidence and capturing the ordinal structure among labels.
The soft-encoded label is defined as: y′n,k = e(−ϕ(yn,rk)) / Σj e(−ϕ(yn,rj)), where ϕ is a distance metric. The resulting Soft-Encoded Cross-Entropy (SCE) loss is: LSCE = −Σn Σk y′n,k log(ŷn,k).
2. Ordinal-Aware Regularization: To address underconfidence that may arise from soft encoding, the authors incorporate a regularization term that explicitly preserves the ordinal structure inherent in the labels.
The regularization term partitions the label space into regions k < yn and k ≥ yn, adjusting logits relative to yn within each region. The penalty function uses a log-barrier-based approach: The logarithmic nature of the log-barrier function enforces strong corrections as the difference r approaches the boundary value −1/t2, indicating a high-risk violation requiring a strong penalty.
The final loss is: LORCU = LSCE + LREG.
The gradient analysis (Table 1 in the paper) shows how the regularization term enforces unimodality and calibration. The authors explain: The difference between neighboring logits, r, reflects both the degree of unimodality and the associated uncertainty in the predicted distribution.
When r ≪ −1/t2, the model's output aligns closely with the ordinal structure, and LORCU applies minimal gradient from LREG. When r ≈ −1/t2, the model encounters a heightened risk of violating unimodality and increased uncertainty,
and the log-barrier function dynamically amplifies the gradient on zk to accommodate this elevated uncertainty.
The paper evaluates ORCU on four datasets: Image Aesthetics (5 ordinal aesthetic score levels, 13,364 images), Adience (8 ordinal age groups, 17,423 images), LIMUC (4-level Mayo Endoscopic Scores for ulcerative colitis, 11,276 images), and Diabetic Retinopathy (5 levels of DR severity, 9,570 images after uniform sampling). Experiments used ResNet-50 pretrained on ImageNet, with additional evaluations on ResNet-34 and ResNet-101.
The evaluation metrics include calibration metrics (SCE, ACE, ECE), ordinal structure metrics (%Unimodality), and classification metrics (Accuracy, QWK, MAE). The authors justify their metric choices: "Unlike Expected Calibration Error (ECE), which uniformly measures calibration across all predictions, SCE computes class-specific calibration, making it more suitable for multi-class tasks with ordinal labels. ACE further refines ECE by dynamically adjusting bin sizes to evenly distribute predictions."
Key results show ORCU achieves state-of-the-art calibration without compromising accuracy. For example, on the Adience dataset, ORCU achieves SCE of 0.4193 compared to 0.8495 for CE baseline and 0.7823 for SORD. On the Image Aesthetics dataset, ORCU achieves SCE of 0.6447 versus 0.7637 for CE and 0.6844 for SORD. The paper states: ORCU strikes an effective balance between calibration and ordinal relationship learning, consistently surpassing existing methods that prioritize only one aspect.
The reliability diagrams (Figure 4) show that CE and LS exhibit overconfidence,
while SORD employs a soft-encoded process to capture ordinal relationships; however, this leads to underconfident predictions and a restricted confidence range.
In contrast, The proposed ORCU effectively mitigates SORD's limitations by incorporating an ordinal-aware regularization term, resulting in more reliable confidence estimates.
The t-SNE visualizations (Figure 5) demonstrate that ORCU generates a well-structured feature distribution with features organized according to label order, demonstrating the effectiveness of its ordinal-aware regularization term,
whereas CE and LS produce dispersed features without an organized ordinal structure.
The softmax output distribution analysis (Figure 6) shows that SORD generates uniform confidence levels, disregarding input-specific uncertainty and failing to adapt dynamically,
while ORCU adjusts confidence levels dynamically based on input characteristics, refining confidence for the correct label according to input uncertainty.
The ablation study (Table 3) demonstrates that Conventional calibration regularization terms that ignore inter-class ordinal relationships disrupt the ordinal structure learned by LSCE and fail to improve calibration,
confirming the necessity of ordinal-aware conditioning to calibrate ordinal tasks effectively.
Among distance metrics, the Squared metric consistently outperformed others.
The authors acknowledge limitations: Although we primarily experimented with computer vision datasets, ORCU has the potential to be employed in various domains such as text, audio, and video.
Future work will optimize the parameters for different domains and validate the method
and address the remaining miscalibration at extreme confidence levels.
The paper concludes: "We proposed ORCU, a novel loss function that improves calibration and unimodality within a unified framework for ordinal classification. By combining soft encoding with an ordinal-aware regularization term, ORCU preserves ordinal relationships while providing reliable confidence estimates. Extensive experiments demonstrate its effectiveness in jointly optimizing calibration and predictive accuracy, addressing critical limitations of existing approaches."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems, along with the resulting capabilities:
-
Implementation: Modify the training loss function to
LORCU = LSCE + LREG, where: -
LSCEuses soft-encoded labels (SORD encoding) with a squared distance metric instead of one-hot labels. -
LREGis an ordinal-aware regularization term that enforces unimodality by penalizing logit differences that violate the expected monotonic increase/decrease around the true label. -
Resulting Capability: The model will produce probability distributions that are both well-calibrated (confidence matches actual accuracy) and unimodal (probabilities peak at the true class and decrease monotonically away from it).
-
Implementation: Integrate the regularization term with a temperature parameter
t = 3.0(determined via validation). The gradient adjustment is: -
For
k < yn(classes below true label): enforcezk zk+1 -
Use a log-barrier penalty when logit differences approach the violation boundary (
r ≤ -1/t2) and a linear penalty otherwise. -
Resulting Capability: The model dynamically adjusts its confidence based on input uncertainty—producing sharper predictions for clear inputs and more conservative, spread-out predictions for ambiguous inputs, while maintaining ordinal consistency.
-
Implementation: During training and evaluation, track Static Calibration Error (SCE) and Adaptive Calibration Error (ACE) instead of relying solely on Expected Calibration Error (ECE). Use these metrics for model selection and early stopping.
-
Resulting Capability: The system will be optimized for per-class calibration, which is critical in imbalanced ordinal datasets (e.g., medical severity levels where rare classes are often miscalibrated).
-
Implementation: After obtaining the softmax output, apply a post-processing step that checks the unimodality condition (
ŷ1 ≤ ŷ2 ≤... ≤ ŷyandŷy ≥ ŷy+1 ≥... ≥ ŷC). If violated, apply a projection algorithm to find the closest unimodal distribution (e.g., using isotonic regression on the logits). -
Resulting Capability: Guarantees that the model never produces contradictory predictions (e.g., predicting class 1 as most likely while giving class 4 the second-highest probability), which is essential for clinical decision support.
-
Implementation: Use the logit difference
r = zk - zk+1(orzk+1 - zkdepending on side) as a proxy for input uncertainty. Whenris near the violation boundary, increase the gradient magnitude on the target class logit to refine predictions; whenris large and positive, apply a fixed correction to prevent overconfidence. -
Resulting Capability: The model will naturally produce higher confidence for easy, clear inputs and lower confidence for ambiguous ones, improving trustworthiness in high-stakes applications.
- In Medical Diagnosis (e.g., Diabetic Retinopathy, Ulcerative Colitis):
-
Provide severity predictions (0–4 or 0–3) with confidence scores that accurately reflect the true likelihood of correctness.
-
Reduce overconfidence on severe cases (which are often rare and critical) by up to 40% in SCE compared to standard CE.
-
Ensure that if the model predicts
Moderate
as most likely, it never assignsSevere
a higher probability thanMild,
maintaining clinical plausibility.
- In Age Estimation (e.g., Adience dataset):
-
Produce age-group predictions (8 ordinal classes) with confidence that is well-calibrated across all age groups, including underrepresented ones (e.g., 60+ years).
-
Achieve 99.94% unimodality in predictions, meaning the model's probability distribution always peaks at the predicted age and decreases monotonically away from it.
- In Image Aesthetics Scoring (5 ordinal levels):
-
Provide reliable confidence estimates for subjective ratings, enabling better human-AI collaboration in content moderation or recommendation systems.
-
Reduce calibration error (SCE) from 0.76 (CE) to 0.64, a 16% improvement, without sacrificing accuracy.
- General Capabilities:
-
Reliable Uncertainty Quantification: The model can distinguish between
I am confident this is class 3
andI am uncertain between class 2 and 3,
which is crucial for triggering human review in automated pipelines. -
Order-Consistent Predictions: The model will never produce logically inconsistent probability distributions, improving interpretability and trust.
-
Robustness to Class Imbalance: With SCE/ACE monitoring and ordinal-aware regularization, the model maintains calibration even when some classes are rare (e.g., only 2% of samples in the highest severity class).
Specific Numerical Improvements (from the paper):
-
SCE reduced by 16–50% across datasets (e.g., 0.85 → 0.42 on Adience)
-
ACE reduced by 16–50% (e.g., 0.84 → 0.41 on Adience)
-
%Unimodality increased to 99.9–100% (from 70–98% with other methods)
-
Accuracy and QWK maintained or improved (e.g., QWK increased from 0.88 to 0.90 on Adience)
Abstract
Recent studies have shown that deep neural networks are not well-calibrated and often produce over-confident predictions. The miscalibration issue primarily stems from using cross-entropy in classifications, which aims to align predicted softmax probabilities with one-hot labels. In ordinal regression tasks, this problem is compounded by an additional challenge: the expectation that softmax probabilities should exhibit unimodal distribution is not met with cross-entropy. The ordinal regression literature has focused on learning orders and overlooked calibration. To address both issues, we propose a novel loss function that introduces ordinal-aware calibration, ensuring that prediction confidence adheres to ordinal relationships between classes. It incorporates soft ordinal encoding and ordinal-aware regularization to enforce both calibration and unimodality. Extensive experiments across four popular ordinal regression benchmarks demonstrate that our approach achieves state-of-the-art calibration without compromising classification accuracy.
Sources
- Non-parametric Uni-modality Constraints for Deep Ordinal Classification
- Deep Residual Learning for Image Recognition
- Unimodal Distributions for Ordinal Regression
- Decoupled Weight Decay Regularization
- Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks