CANDOR: Chance-Calibrated Neighborhood Discordance in Frozen Encoders for Medical Imaging
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "CANDOR: Chance-Calibrated Discordance in Frozen Foundation Encoders".
Jane: The paper was written by Soroosh Tayebi Arasteh, Sven Nebelung and Daniel Truhn from Lab for AI in Medicine, RWTH Aachen University and Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Summary: Tom: So, we’ve established that CANDOR is a powerful diagnostic tool for measuring how far a foundation model might be from the truth in a given context. Now, let's talk about what the paper summarizes about its core methodology and why it’s so effective.
Jane: The summary highlights that while general-purpose models are trained on huge amounts of data, they often fail to capture fine-grained details like specific medical findings because their architecture simply isn't designed to hold that detail.
Meng: This is a major concern for deployment; if we see a finding placed closer to its opposite in the feature space, it means the model lacks the necessary internal representation for a high-confidence diagnosis.
Lu: I found it fascinating how they quantify this as "discordance," which suggests that even though the model might be trained on millions of images, its ability to relate two specific points is what's truly limiting its effectiveness.
Lalam: For AI culture, this means we are seeing a clear gap between how broad the training data is and how deep the actual understanding is for highly specialized human knowledge systems like radiology.
Jane: The core idea isn't that the model is bad, but that we need a better way to *measure* how far off its general knowledge was from the specific, narrow context of what it’s being used for right now.
Tom: So, to circle back to that idea of "discordance," if I try to simplify it, are they suggesting there's a mathematical distance between the model's internal state and the truth of a given premise?
Jane: That's a great way of putting it, Tom; it gives us a quantifiable measure that goes beyond simple input/output checking; it looks at the internal probability space itself.
Meng: Practically speaking, this means we could build an automated pipeline that runs an input through CANDOR first. If the discordance score is too high, instead of trusting the output, we could trigger a warning or request human review immediately.
Lu: And that capability is revolutionary for building reliable autonomous agents! We're moving from simply asking "Did it work?" to "How sure are we that it worked, and why?"
Lalam: This ability to quantify uncertainty in the foundational knowledge improves the overall human-AI trust relationship because it makes the model's internal reasoning process visible to us, even if only statistically.
Tom: Wow, so we’re talking about building explicit failure modes into our AI systems. It sounds like a critical shift from simply maximizing performance to maximizing *reliability* and *understandability*.
Jane: It does. And it sets the stage perfectly for discussing how they actually improve upon these existing measurements.
Improvements: Tom: We've established that CANDOR is a powerful diagnostic tool for measuring how far a foundation model might be from the truth in a given context. Now, let's talk about what the paper suggests for *improving* this method and making it more reliable.
Jane: The authors propose specific refinements to fix issues with simply calculating discordance, moving beyond just suggesting a new metric to improving the structural integrity of how we interpret the signal.
Lu: What I took away from reading this section is that the original methods might be too sensitive to minor fluctuations in input data, which makes them impractical for reliable use because the score jumps around wildly without reflecting true conceptual shifts.
Meng: If it's too sensitive, then deploying it would be impossible because we'd have to filter out almost all the readings as noise; we need robust, stable indicators of a knowledge gap.
Lalam: For culture and adoption, stability is everything; if a system gives erratic warnings about its own shortcomings, users will quickly lose faith in its ability to help them with critical tasks. The improvements must build durable trust signals.
Jane: Exactly, Lalam. They address this by refining the mathematical framework so that a high discordance score really means something significant about the the model's foundational knowledge state, rather than just being a statistical anomaly.
Tom: So, they’re essentially making the diagnostic tool itself more robust—less prone to false positives or meaningless fluctuations? Is that right?
Jane: Pretty much, Tom. They're stabilizing the calculation so that even with minor changes in the data, the signal remains consistent and reliable across multiple uses.
Lu: I think one of the most exciting improvements is how they link this measurement not just to input data, but to specific conceptual subspaces within the model itself, allowing us to identify exactly where a failure is occurring.
Meng: From an engineering standpoint, that targeted intervention capability changes everything; instead of needing a full re-training cycle when discordance is found, we might be able to target local adjustments based on the diagnosed gap.
Lalam: Pinpointing where a knowledge deficit lives allows for more informed decisions about how we should interact with AI systems, improving our overall cooperation with these powerful models.
Tom: This is how we transition from knowing a model is struggling to building strategies that allow us to recover and improve the system's performance.
Paper discussion segment 3: Tom: We’ve seen how CANDOR works and what refinements are suggested, but let's dig deeper into the findings—specifically, how does this concept of "weakness" relate to the actual capacity of a model to learn?
Jane: The authors use Proposition one to show that when a model has a discordant twin—a finding placed near its opposite—that limits the maximum possible margin any subsequent head can achieve.
Meng: That’s profound because it means the failure is purely structural, not about lack of data; we have a hard mathematical limit on how much separation the AI can possibly create between those two concepts.
Lu: It’s interesting that even though this geometric cap exists, some parts of the representation are still functioning correctly, showing that the overall "weakness" is highly specific and localized to certain types of findings.
Lalam: This highlights a nuance in AI failure: we can't just say "the model failed," we have to say *how* it failed, and CANDOR gives us that precise map of the frozen knowledge.
Tom: The paper also raises the question of whether this failure is information-related or something else entirely—what does the authors say about that?
Jane: They argue that while it looks like a lack of information, there's an open door: we could potentially select a specific encoder for each image, but that selection process isn't constrained by any single geometric rule.
Meng: That means the solution to fixing this problem isn't just in the model itself, but in having a system that can choose the best tool for the job, which is an incredible engineering challenge.
Lu: We are seeing how we need to distinguish between a fundamental conceptual failure and merely trying to find a way around it by leveraging different parts of its internal structure.
Lalam: The ability to even see these specific failures allows us to move toward designing truly self-aware systems, which is the next big step in AI development.
Conclusion: Tom: So, we’ve covered how CANDOR works, what improvements are needed for a reliable tool, and where its limits lie. What's the final word on what this means for a big clinical system using these foundation models?
Jane: The core message is that foundation models aren't just randomly poor; they are often weakly encoding specific, fine-grained details of the world—like subtle lesions or rare diseases.
Lu: That’s such a powerful shift in perspective, moving from "it failed" to "it failed in this specific way," which really helps us understand the nature of representation itself.
Meng: Practically, this means we can stop assuming general AI is ready for specialized tasks and start designing systems that know exactly where their own knowledge gaps are.
Lalam: It tells us that recognizing failure isn’t a signal of weakness, but a map showing precisely where the model needs to be trustworthy.
Tom: A map of its own limitations—that’s a huge concept for transparency in AI decision-making.
Jane: It lets us move away from simply chasing high overall performance and focusing on *why* by targeting specific types of failure, like those rare clinical findings.
Lu: We are seeing the geometry of knowledge, really; it's not just about what’s wrong with the output but how its internal space is organized.
Meng: The engineer in me sees a system that is far more reliable because we know exactly when to trigger a human intervention based on this measurable collapse.
Lalam: By understanding this "Chance-Calibrated Discordance," we are improving the culture of trust, allowing for AI that can admit its shortcomings before failing catastrophically.
Tom: It’s clear that the authors, with "CANDOR: Chance-Calibrated Discordance in Frozen Foundation Encoders," have given us a way to see inside the black box and understand its structural weaknesses.
Jane: And it shows that even when using general-purpose foundation models, we can find ways to measure their specific weaknesses without needing a model for the task.
Lu: It’s proof that understanding how AI fails is fundamental, even if the failure isn' isn't tied to one specific architecture or a single data point.
Meng: The practical impact of "CANDOR: Chance-Calibrated Discordance in Frozen Foundation Encoders" is massive because it provides the framework for building self-aware systems, which is exactly what we need moving forward.
Lalam: It gives us a new vocabulary for discussing AI reliability that truly respects its current limitations while striving for the future.
Lab for AI in Medicine, RWTH Aachen University · Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen
cs.LG, cs.AI, cs.CL, cs.CV
Submitted: 2026-07-20
Updated: 2026-09-20
Code: https://github.com/tayebiarasteh/candor
Importance score: 89/100
The gist: This paper introduces CANDOR (Chance-Calibrated Neighborhood Discordance), a new framework for measuring how well frozen foundation encoders represent fine-grained clinical findings.
Key concepts
- CANDOR: Chance-Calibrated Discordance
- This is a diagnostic tool that measures the mathematical distance between a foundation model's internal knowledge state and the truth of a given premise. It provides a quantifiable measure of how far off the model might be from accuracy, going beyond simple input/output checks.
- Discordance
- Discordance is the core metric used by CANDOR to quantify uncertainty. It suggests there is a measurable mathematical distance between the model's internal representation and a specific truth. A high score indicates a significant gap in foundational knowledge for that context.
- Frozen Foundation Encoders
- These are general-purpose models trained on vast amounts of data but used for specialized tasks, like reading medical images. The concept highlights that while the models are broad, they may fail to capture the fine-grained details needed for high-confidence diagnoses.
Terminology
Summary
This paper introduces CANDOR (Chance-Calibrated Neighborhood Discordance), a new framework for measuring how well frozen foundation encoders represent fine-grained clinical findings. It addresses a critical flaw in current evaluation methods where encoders are judged solely by how well a lightweight head performs on their features, which can mask geometric failures where positive cases sit closer to negative neighbors. By providing a chance-calibrated measure, CANDOR reveals that many state-of-the-art encoders are not blind
to medical findings but are significantly weak
in their representation geometry.
The Problem with Current Metrics
Current evaluation of foundation models relies on the Area Under the Receiver Operating Characteristic curve (AUROC) of a lightweight head. The authors argue this cannot see the failure that governs the risk,
as a high AUROC can still coexist with a large minority of positive cases sitting nearer to their opposite label in feature space. Existing geometric measures attempt to address this but are confounded by prevalence
; when reference banks are unequal in size, the opposite-label neighbor wins on density rather than geometry, causing an uncorrected estimator to report the most collapse for the rarest findings.
The CANDOR Framework
CANDOR fixes these issues by using equal-size reference banks of positive and negative samples within a matched acquisition context (such as same hospital and projection). This design choice makes the label groups exchangeable under relabeling,
which fixes the chance level at exactly one half. The core operator measures:
The share of positives an encoder places nearer to the opposite label than their own kind.
Because the chance level is fixed at 50%, CANDOR can be read before any head is trained, allowing researchers to flag findings that a frozen encoder supports poorly.
Key Findings and Results
Across 22 encoders and over 605,443 images, the study reveals a map of the frozen regime.
While some tasks like bird species recognition show very low discordance (e.g., DINOv3-L at 4.5%), medical tasks show significant weakness:
Chest findings at a median 42.1 and glaucoma at chance.
Even the strongest model, RAD-DINO, which achieves an 84.5 AUROC on pneumothorax, still places 18.4% of those positives nearer an opposite-label film than its own kind. The authors prove that this discordance caps the normalized margin of any Lipschitz head,
meaning no amount of architectural scaling or training can overcome the underlying geometric deficit in the encoder's representation.
Mechanisms and Correlations
The paper investigates what drives this discordance through three specific instruments:
The blind set:
Collecting positives that an encoder confuses with their opposite. The authors find that failure is often shared across different encoders in medical domains.
Erasure retention:
Measuring how much a representation moves when the evidence region is occluded. The study finds erasure retention is associated with collapse,
showing a Spearman correlation of 0.855; encoders that barely move when evidence is removed are the ones that place positives near their opposites.
Neighborhood impurity:
A label-free score intended to flag discordant twins, though the authors find it does not work
as it fails to outperform standard confidence metrics like maximum softmax probability.
Ultimately, the paper concludes that while encoders are not broadly blind, they fail on specific types of evidence where the representation is least sensitive to the clinical finding. The deficit is selection, not information,
meaning that while the information exists in a panel of models, reaching it requires a rule that can select the correct encoder per image.
Improvements for AI systems
To improve AI systems based on the CANDOR framework, I would implement the following specific technical improvements:
-
Implement a
CANDOR-based Geometric Uncertainty Filter
for frozen foundation models. -
Integrate
Erasure Retention Scoring
as a secondary validation layer during inference.
By implementing these, the improved AI system can:
-
Identify and flag
Discordant Twins
: The system will detect when a medical image is geometrically closer to the opposite clinical label (e.g., a pneumothorax film sitting near healthy films) within its specific acquisition context (same hospital/projection). This allows the system to preemptively flag high-risk cases where the encoder's feature space is known to beblind
or weak, even if the classification head reports high confidence. -
Execute
Evidence-Sensitivity Verification
: By performing a real-time erasure test (masking out annotated clinical evidence), the system can determine if the model's representation actually shifts in response to the finding. If the representation remains static despite removing key diagnostic features, the system will trigger anInformation Deficit
alert, signaling that the model is relying on spurious correlations rather than actual clinical evidence. -
Enable
Per-Image Encoder Selection
: Instead of relying on a single general-purpose encoder, the system can use a non-Lipschitz selection rule to choose between multiple foundation models for every specific image. This bypasses theLipschitz margin cap
described in the paper, allowing the system to achieve higher accuracy by dynamically switching to whichever model's geometry is most robust for that specific patient's visual features.
Sources
- Understanding intermediate layers using linear classifier probes
- Vision-language models for chest radiography do not always need the image
- Agentic retrieval-augmented reasoning reshapes collective reliability under model variability in radiology question answering
- The strength of clinical evidence is recoverable from language model representations but not from their stated grades
- Resolution scaling governs DINOv3 transfer performance in chest radiograph classification
- FairVision: Equitable Deep Learning for Eye Disease Screening via Fair Identity Scaling
- Phikon-v2, A large and public feature extractor for biomarker prediction
- DINOv3
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks