Subtype Robustness Is Not Just Accuracy: Calibration Under Unseen Subtype Shift

arXiv:2608.00928 · cs.LG · Submitted 2026-08-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Subtype Robustness Is Not Just Accuracy: Calibration Under Unseen Subtype Shift".

Jane: The paper was written by Hanyu Su, Carlota Julbe i Juanola and Yibo Hu from Illinois Institute of Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, folks. Today we're digging into a paper with a mouthful of a title: "Subtype Robustness Is Not Just Accuracy: Calibration Under Unseen Subtype Shift." Jane, I gotta say, even the title is telling us something important.

Jane: It really is, Tom. And I love it because it's basically saying, "Hey, we've been measuring the wrong thing." For years, when we test whether a model is robust to new data, we just look at accuracy. Did it get the answer right or wrong? And this paper says that's not enough.

Tom: Right, because a model can be accurate and still be totally unreliable in a sneaky way. Let me give the listeners the setup. Imagine you train a model to recognize vehicles and animals. It sees cars, airplanes, cats, and birds. Now you test it on a ship and a deer. Those are new subtypes, but they still belong to the known supertypes of vehicle and animal.

Jane: And that's the key distinction. This isn't out-of-distribution detection, where the input is totally alien. The ship is still a vehicle. The deer is still an animal. The model should be able to handle that, right? Well, it turns out it handles it in a really dangerous way.

Tom: Dangerous how? That's the million-dollar question.

Jane: So the paper shows that when models see these unseen subtypes, their accuracy drops dramatically. We're talking thirty to forty percentage points on some datasets. But here's the kicker—their confidence barely drops at all. The model is just as sure of itself, even when it's wrong.

Tom: So you've got a model that's confidently wrong. It thinks it knows the answer, but it doesn't. And that's a calibration failure. The model's confidence no longer matches its actual chance of being correct.

Jane: Exactly. And the authors call this "silent overconfidence." The model doesn't know it's in trouble, and neither does the person relying on it. It's not like when you show a model a blurry, corrupted image and it gets less sure of itself. That's a normal reaction. This is different. The image looks perfectly fine, it fits the taxonomy, but the model is just guessing while acting like it knows.

Tom: And that's why this paper matters so much. It's not just an academic curiosity. Think about a species classifier used in the wild. It's trained on known species, but it will always encounter new ones. If it's confidently wrong about those, conservationists might make bad decisions based on bad data.

Jane: Or think about medical imaging, where new variants of diseases appear all the time. The coarse label might be "skin lesion," but the specific subtype is new. If the model is overconfident on that, a doctor might trust a wrong diagnosis.

Tom: So the title is really a warning. Subtype robustness isn't just about keeping accuracy high. It's about keeping the model honest about what it knows. And this paper is the first to really nail that down across so many datasets and architectures.

Jane: And we're just getting started. We've got the team here to break down how they proved this, what it means for fixing models, and whether the standard tools can save us. Stick around.

Summary: Tom: Alright, we're back with "Subtype Robustness Is Not Just Accuracy: Calibration Under Unseen Subtype Shift." And Jane, we set the stage. Now let's get into the meat of what they actually did. They ran this across ImageNet, BREEDS, iNaturalist, CIFAR-one hundred and five different architectures.

Jane: And the results are remarkably consistent, Tom. On every single model, on every dataset, they saw the same dissociation. Accuracy falls off a cliff, but confidence barely moves. Let me pull up a specific number. On ImageNet-twenty-five with a ResNet-fifty accuracy drops by about thirty points on unseen subtypes. Confidence only drops by about eleven points.

Tom: So the gap between those two numbers is where the overconfidence lives. And that gap shows up as a massive jump in calibration error. The expected calibration error, or ECE, goes from about two point five points on seen subtypes to over twenty-one points on unseen ones. That's a huge deal.

Jane: It is. And Lu, you've been nodding along. What's your take on this from a research perspective?

Lu: I think the most striking part is that they didn't just show this happens. They showed it's specific to subtype shift. They ran a control experiment using generic image corruptions—things like blur, noise, compression artifacts. They tuned the corruption severity so that the accuracy loss was exactly matched to the subtype shift.

Meng: So they made the two conditions equally hard in terms of accuracy, right?

Lu: Precisely, Meng. And under matched accuracy loss, the generic corruption caused a much bigger drop in confidence. On ImageNet-twenty-five corruption cost about eighteen point six points of confidence, while subtype shift only cost ten point eight. The model reacts to visible degradation, but it stays blind to in-taxonomy novelty.

Meng: That's a really clean experimental design. It rules out the explanation that any accuracy loss just naturally makes a model less confident. The model is specifically failing to notice that the input is novel in a semantic way.

Tom: And that's the scary part. The image looks normal. The label is still valid. There's no reason for the model to doubt itself, so it doesn't. It just confidently makes mistakes.

Jane: And they also checked whether this is just an aggregate problem or if it affects per-sample reliability. They looked at failure-detection AUROC, which measures whether confidence can tell correct from incorrect predictions. That dropped by zero point one five to zero point one eight on the big benchmarks. So it's not just that the average confidence is too high. The ranking of which predictions to trust is also broken.

Lu: Right. And that's a deeper problem. You could fix the average confidence with a simple rescaling, but you can't fix the ranking. If the model is confident about the wrong things and unsure about the right things, no monotone transformation of confidence will save you.

Meng: So we've got a model that's overconfident, and the overconfidence is tied to the semantic novelty, not the visual appearance. What's the mechanism? Why does this happen?

Lu: The paper points to a nice explanation. When you train on only a few subtypes per supertype, the model learns to compress those subtypes into a single feature region. An unseen subtype from the same supertype lands right in that region, so the model treats it as familiar. It's not that the features are wrong—it's that they're too coarse to notice the novelty.

Tom: And that's why the number of seen subtypes matters. The paper shows that supertypes with very few seen subtypes, like "fungus" with just one, collapse completely. But "dog" with one hundred seventeen seen subtypes stays well-calibrated even on unseen dogs. The model needs diversity to learn that the supertype is a broad category, not a specific look.

Jane: So the fix isn't just about training longer or harder. It's about the structure of the training data itself. And that's a profound insight for anyone building these systems. But we still have to ask—can we fix this after the fact? That's coming up next.

Improvements: Tom: Welcome back. We're still on "Subtype Robustness Is Not Just Accuracy: Calibration Under Unseen Subtype Shift." And Jane, we've established the problem. The model is silently overconfident on unseen subtypes. Now the big question—can we fix it?

Jane: That's exactly what the paper tackles next, Tom. And the news isn't great. They tried the standard post-hoc remedies, and they all fall short. First up, recalibration. You take a held-out set of seen subtypes, fit a temperature scaling parameter, and apply it to the unseen subtypes.

Meng: So you're tuning the model's confidence on data it has seen, hoping it transfers to data it hasn't. That sounds like it could go wrong.

Jane: It does, Meng. Temperature scaling does lower the unseen calibration error. On ImageNet-twenty-five it goes from twenty-one point seven down to seventeen point eight. That sounds like progress, right? But the seen calibration error drops to almost nothing, one point four points. So the gap between seen and unseen actually remains huge.

Lu: And that's the key point. The recalibration fixes the split it was tuned on, but it doesn't generalize. The model is still overconfident on unseen subtypes, just slightly less so. The paper calls this "narrowing but not closing" the gap.

Tom: So it's like putting a band-aid on a wound that needs stitches. It helps a little, but the underlying problem is still there.

Meng: What about the other approach? Instead of fixing the confidence, can we at least detect when the model is seeing an unseen subtype? If we know it's novel, we can flag it for a human.

Jane: That's the second remedy they tested. They used three standard out-of-distribution detection scores: maximum softmax probability, energy score, and Mahalanobis distance. And the results are pretty disappointing.

Tom: The output-based scores, MSP and energy, only hit about zero point seven two AUROC on the big benchmarks. That's above chance, but it's nowhere near usable. At a threshold that flags five percent of seen inputs, they only catch twelve to sixteen percent of unseen ones.

Lu: And the Mahalanobis distance is even worse. It's actually below chance, around zero point four five. That means the model ranks unseen subtypes as slightly less novel than seen ones. The feature space has collapsed the subtypes together so thoroughly that the detector can't tell them apart.

Meng: So the model's representation is the problem. It's not just the classifier head. The features themselves don't preserve enough subtype information to detect novelty.

Jane: Exactly. And that's why this is so hard to fix post-hoc. You can't detect what the model can't represent. The information is just gone.

Tom: But wait, they did find one thing that helps. The paper shows that pretraining on ImageNet-1K before fine-tuning on the coarse labels reduces the gap. The calibration gap drops from nineteen point two to twelve point eight on ImageNet-twenty-five. It doesn't close it, but it helps.

Lu: That makes sense. A pretrained backbone has richer features that preserve more subtype distinctions. It's not trained to throw away that information. So when it encounters an unseen subtype, it has more to work with.

Meng: So the lesson is that the fix has to happen at training time, not after. You need to either use pretrained features or deliberately train for subtype diversity.

Jane: And that's the practical takeaway. If you're building a system that will encounter new subtypes, you can't just bolt on a calibration step and hope for the best. You have to build the robustness in from the start.

Tom: Alright, so we've got the problem, the evidence, and the failed remedies. But there's one more thing—the paper found a signal that predicts which supertypes will collapse. That's coming up in our final segment.

Conclusion: Tom: We're wrapping up our discussion of "Subtype Robustness Is Not Just Accuracy: Calibration Under Unseen Subtype Shift." And Jane, we've covered a lot. Let's bring it home with that last finding.

Jane: Right, Tom. The paper found that the number of seen subtypes per supertype is a strong predictor of overconfidence. On ImageNet-twenty-five supertypes with few seen subtypes, like "big cat" with only three, collapse completely. Their accuracy drops from ninety-six percent to five percent, and their calibration error skyrockets to eighty-one points.

Lu: But "dog" with one hundred seventeen seen subtypes stays calibrated. The Spearman correlation between seen-subtype count and overconfidence is negative zero point eight two. That's a very strong relationship.

Meng: So it's a diversity problem. If you train on enough subtypes, the model learns that the supertype is a broad category. If you train on too few, it overfits to the specific look of those subtypes.

Jane: And they even ran a controlled experiment to prove it's causal. They held the image budget fixed and varied only the number of seen subtypes from one to four. The calibration gap dropped monotonically from thirty-one to twenty-three points. More subtypes, less overconfidence. Period.

Tom: So the message is clear. If you're building a model for a hierarchical taxonomy, you need to audit it for calibration, not just accuracy. And you need to be especially worried about the supertypes that are sparsely sampled in your training data.

Lu: And the paper's final point is that this isn't a problem that post-hoc tools can solve. Recalibration narrows the gap. Detection scores barely work. The fix has to be in the training data or the training objective.

Meng: For me, the practical implication is that we need to report confidence reliability alongside accuracy whenever we claim subtype robustness. Otherwise, we're hiding the failure.

Jane: And that's the lasting impact of this paper. It gives us a new metric to watch, a new failure mode to design for, and a clear warning about the limits of our current tools.

Tom: Well said, Jane. And with that, we're saying goodbye to "Subtype Robustness Is Not Just Accuracy: Calibration Under Unseen Subtype Shift." A fantastic piece of work from the team at Illinois Institute of Technology. Thanks to Lu, Meng, and our in-house LLM Lalam for joining the conversation.

Lalam: And from a cultural perspective, this paper reminds us that as we deploy AI in biodiversity monitoring, medicine, and other high-stakes fields, we need systems that know when they don't know. That humility is what builds trust.

Tom: Couldn't agree more. That's all for today, folks. Join us next time when we'll be diving into another fresh arXiv paper. Until then, keep questioning what your models are telling you.

Hanyu Su, Carlota Julbe i Juanola, Yibo Hu

Illinois Institute of Technology

cs.LG

Submitted: 2026-08-14

Updated: 2026-08-17

Comments: 15 pages, 6 figures, 12 tables

Code: https://github.com/yibo-hu-lab/subtype-robustness

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 74/100

The gist: This paper investigates subtype robustness in deep learning classification systems: whether a model keeps the correct coarse prediction when test examples come from fine-grained subtypes absent from

Key concepts

Subtype Shift
Subtype shift occurs when a model trained on specific examples encounters a new subtype that belongs to the same general category or 'supertype.' This is different from completely foreign data, but the model fails to recognize this semantic novelty.
Calibration Failure
A calibration failure happens when a model's confidence in its prediction does not match its actual probability of being correct. The paper describes this as 'silent overconfidence,' where the model is highly certain of a wrong answer.
Subtype Robustness
This concept refers to how well a model performs on unseen subtypes within a hierarchical taxonomy. The episode argues that robustness requires more than just high accuracy; it demands that the model is honest about its confidence.

Terminology

Summary

This paper investigates subtype robustness in deep learning classification systems: whether a model keeps the correct coarse prediction when test examples come from fine-grained subtypes absent from training but still inside a known coarse category. The authors argue that prior work has studied this almost entirely through accuracy, and they ask whether the model also stays calibrated. They present the first systematic study of the question across ImageNet, BREEDS, iNaturalist and CIFAR-100 with five architectures.

The setting is defined as follows: test examples may come from fine-grained subtypes absent from training, while their correct coarse labels remain within the known taxonomy. This is distinct from classical OOD detection because the input still carries a valid coarse label, so the failure is a wrong coarse prediction.

The paper formalizes the setting with a hierarchical label structure. Multiclass classification maps instances to K labels, partitioned into J mutually exclusive supertypes. They define:

  • Seen subtypes S: subtypes used in training

  • Unseen subtypes U: held-out subtypes under the same supertypes

  • Every supertype is represented among seen subtypes: g(S) = g(U) = J

All models are trained using only examples whose subtype labels belong to S, and predict the supertype. The key constraint: no examples from U are used for representation learning, classifier training, calibration, or model selection. The training procedure is direct-coarse: a single network optimized end-to-end on the supertype label from a random initialization.

The paper defines three matched gaps:

  • ∆acc = accuracy on seen − accuracy on unseen

  • ∆conf = mean confidence on seen − mean confidence on unseen

  • ∆cal = ECE on unseen − ECE on seen

They use expected calibration error (ECE) over 15 equal-width bins, corroborated with equal-mass binning and two binning-free proper scores (NLL and Brier score). They also report per-sample discrimination via failure-detection AUROC and risk–coverage AURC.

Five benchmarks are used:

  • ImageNet-25 (IN-25): 25 broad categories as supertypes, 383 seen subtypes, 3,872 unseen subtypes

  • BREEDS: Living-17 (17 supertypes) and Nonliving-26 (26 supertypes), each split into two seen and two unseen subtypes per supertype

  • iNaturalist (iNat-25): 25 orders as supertypes, 3,427 seen and 3,427 unseen species

  • CIFAR-100: 20 supertypes, 60 seen and 40 unseen fine classes

Five architectures are compared: AlexNet, ResNet-18, ResNet-50, ViT-B/16, and ConvNeXt-T (CIFAR-100 uses only AlexNet and ResNet-18).

The paper reports: "Coarse accuracy drops by 29–40 points on CIFAR-100, IN-25 and BREEDS and by a milder 7–13 on iNat-25, but confidence drops by only 2–16. On unseen subtypes accuracy lands at 30–70% while confidence stays at 64–90%. ECE therefore rises from 2–22 points on seen subtypes to 9–46 on unseen, so miscalibration increases by 5–28 points on every model."

The dissociation is summarized: confidence barely follows, leaving the model overconfident where it has become less accurate. The paper calls this silent overconfidence.

Per-sample discrimination also degrades: On unseen subtypes the failure-detection AUROC falls on every dataset, by 0.15–0.18 on CIFAR-100, IN-25 and BREEDS and by 0.05 on iNat-25, and the risk–coverage AURC rises 2.3× to 18×.

The paper compares against ImageNet-C corruptions applied to seen-subtype images, matching accuracy loss: At matched accuracy loss, generic corruption costs 17.3 points of confidence on average against 9.9 for subtype shift (18.6 vs. 10.8 on IN-25, 19.0 vs. 8.9 on Liv-17, 23.1 vs. 16.2 on NL-26, 8.3 vs. 3.8 on iNat-25).

The interpretation: The model loses the same accuracy either way, yet on every dataset it gives up less confidence to the subtype shift: it reacts to visible degradation and stays blind to in-taxonomy novelty.

Temperature scaling and vector scaling are tuned on seen subtypes and applied to unseen subtypes. The paper reports: Temperature scaling lowers the unseen ECE on every dataset, by 18–57% (IN-25 21.7 → 17.8, iNat-25 19.3 → 8.3). However, what the calibrator fixes is the split it was tuned on. The calibrated unseen ECE stays far above the calibrated seen ECE. Vector scaling improves the seen split most and transfers worst to unseen subtypes. The conclusion: Recalibration has been reported to transfer across hierarchy levels for vision-language models. That transfer does not hold under discriminative subtype shift.

Three detection scores are evaluated: MSP, energy score, and Mahalanobis distance. The results: The two output-based scores, MSP and energy, reach only ∼0.70 on the three larger benchmarks and 0.56 on iNat-25. The Mahalanobis distance sits just below chance on every dataset (0.45–0.47), ranking unseen subtypes as slightly less novel than seen ones. The paper explains: A network trained only on the coarse label collapses same-supertype subtypes together, so an unseen subtype ends up looking like the seen ones in feature space.

The paper finds that the collapse is heterogeneous across supertypes: some coarse classes keep near-perfect calibration on unseen subtypes, others collapse completely. The predictor that tracks the collapse is the number of seen subtypes per coarse class, nj: supertypes with few seen subtypes are the most overconfident on unseen ones while richly-sampled ones (e.g. dog, 117 subtypes) stay calibrated (Figure 5; Spearman ρ = −0.82).

A controlled k-sweep supports a causal reading: holding the per-class image budget fixed and varying only the seen-subtype count shrinks ∆cal monotonically from 31 (k=1) to 23 (k=4).

The paper reports two controlled variations:

  • Protocol (one-stage direct-coarse vs. two-stage fine-pretrained-frozen): "Protocol is neutral... switching one-stage direct-coarse for two-stage fine-pretrained-frozen leaves the unseen calibration gap essentially unchanged (∆∆cal < 6)."

  • Initialization (from-scratch vs. ImageNet-pretrained): Pretraining reduces but does not remove the gap: on IN-25 an ImageNet-1K-pretrained frozen backbone lowers ∆cal (19.2 → 12.8) without closing it.

The paper's practical upshot: subtype robustness must be audited with calibration, since accuracy alone hides the failure. The two failure modes are complementary: an aggregate miscalibration that a single temperature could repair only with oracle knowledge of which inputs are unseen, and a per-sample discrimination loss that no monotone rescaling can touch.

The paper concludes: Preserving coarse-label accuracy is not enough if confidence no longer reflects reliability, so subtype robustness should be treated as a joint requirement on prediction and confidence.

Future directions mentioned include dense prediction, language and tabular models, training-time objectives (evidential learning, outlier exposure, subtype-diversity augmentation), and detectors designed for in-taxonomy novelty.

Table 1 (main results, selected rows for ResNet-50):

  • IN-25: seen accuracy 91.0, unseen 60.7, ∆acc 30.3; seen confidence 93.4, unseen 82.3, ∆conf 11.2; seen ECE 2.5, unseen 21.7, ∆cal 19.2

  • Liv-17: seen accuracy 89.2, unseen 56.8, ∆acc 32.4; seen confidence 93.4, unseen 84.4, ∆conf 8.9; seen ECE 4.5, unseen 27.6, ∆cal 23.1

  • NL-26: seen accuracy 84.0, unseen 45.3, ∆acc 38.6; seen confidence 88.1, unseen 71.8, ∆conf 16.2; seen ECE 4.4, unseen 26.5, ∆cal 22.1

  • iNat-25: seen accuracy 76.2, unseen 63.6, ∆acc 12.6; seen confidence 86.6, unseen 82.9, ∆conf 3.6; seen ECE 10.3, unseen 19.3, ∆cal 9.0

Table 2 (post-hoc calibration, ResNet-50): Temperature scaling on IN-25 reduces unseen ECE from 21.7 to 17.8, leaving ∆cal at 16.4; on iNat-25 from 19.3 to 8.3, leaving ∆cal at 6.9.

Table 3 (novelty detection AUROC, ResNet-50): MSP 0.723 (IN-25), 0.723 (Liv-17), 0.740 (NL-26), 0.563 (iNat-25); energy similar; Mahalanobis 0.453–0.473.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:


Current system flaw: Models are evaluated on coarse-label accuracy alone, which hides systematic overconfidence on unseen subtypes.

Improvement: Add a three-part evaluation protocol that reports:

  • ∆acc (seen minus unseen accuracy)

  • ∆conf (seen minus unseen confidence)

  • ∆cal (unseen minus seen ECE)

What the improved system can do: When deployed on new fine-grained variants of known categories, the system will explicitly flag cases where confidence stays high while accuracy drops—preventing silent overconfidence that could cause costly misclassifications in production.

Current system flaw: The model reports high confidence on unseen subtypes even when it is wrong, with no internal signal to warn the operator.

Current system flaw: Supertypes with few seen subtypes (e.g., big cat with 3 seen subtypes) show ECE rising from 2 to 81 on unseen subtypes, while richly-sampled ones (e.g., dog with 117 subtypes) stay calibrated.

Current system flaw: Standard OOD scores (MSP, energy, Mahalanobis) perform poorly at detecting unseen subtypes (AUROC 0.45–0.74), because the model treats them as normal inputs.

Current system flaw: A single temperature tuned on seen subtypes leaves a 7–17 point calibration gap on unseen subtypes.

Specifically, set T 2 = T 1 times (1 + alpha / sqrt n j), where alpha is learned from a small calibration set and n j is the seen-subtype count per supertype.

Current system flaw: The risk–coverage AURC rises 2.3×–18× on unseen subtypes, meaning the same confidence threshold lets through far more errors.

Current system flaw: Models trained on few seen subtypes become overconfident on unseen ones because they collapse same-supertype features together.

The improved AI system will:

  1. Detect when it is operating on unseen subtypes (via the novelty detector)

  2. Report calibrated confidence even on those subtypes (via subtype-aware temperature)

  3. Abstain when its confidence is unreliable (via risk-coverage policy)

  4. Warn operators when silent overconfidence is likely (via the early-warning trigger)

  5. Maintain calibration across diverse subtype distributions (via the diversity regularizer)

These changes directly address the paper's core finding: subtype robustness requires calibration, not just accuracy. The improved system will be trustworthy in production settings where new fine-grained variants of known categories are expected.

Abstract

Subtype robustness asks whether a model keeps the correct coarse prediction when test examples come from fine-grained subtypes absent from training but still inside a known coarse category. Prior work studies this almost entirely through accuracy. We ask whether the model also stays calibrated. We present the first systematic study of the question across ImageNet, BREEDS, iNaturalist and CIFAR-100 with five architectures. Calibration breaks down on unseen subtypes, where accuracy drops while confidence barely follows, leaving the model systematically overconfident exactly where it has become less accurate. At matched accuracy loss, generic image corruption causes a much larger drop in confidence, so the effect is not a general consequence of losing accuracy. The model reacts to visible degradation but not to in-taxonomy novelty. Recalibration tuned on seen subtypes narrows the gap but does not close it, and out-of-distribution scores flag the affected inputs only weakly. Subtype robustness should therefore be evaluated through calibration, not accuracy alone.

Sources

Related papers