Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning
summary
The gist
The gist The Similarity-as-Evidence (SaE) framework calibrates text–image similarities by introducing a Similarity Evidence Head (SEH), which reinterprets the similarity vector as evidence and
In short
The Similarity-as-Evidence (SaE) framework recalibrates overconfident Vision-Language Models for medical active learning by treating similarity scores as evidence rather than deterministic predictions. It introduces a Similarity Evidence Head to quantify model certainty, decomposes uncertainty into knowledge gaps (vacuity) and decision conflicts (dissonance), and uses a dual-factor acquisition strategy to select the most informative samples, leading to improved accuracy and calibration.
Key concepts
- Similarity Evidence Head (SEH)
- This component maps raw text-image similarity scores into Dirichlet evidence parameters. It is trained with a loss function that balances classification performance against the model's intrinsic certainty, allowing the framework to quantify how much evidence the VLM has for each prediction.
- Uncertainty Decomposition
- The framework breaks down uncertainty into two parts: vacuity, which represents knowledge gaps or missing information about rare phenotypes, and dissonance, which measures conflicts between competing class hypotheses. This decomposition provides a more nuanced understanding of why a model is uncertain.
- Dual-Factor Acquisition Strategy
- This strategy selects samples based on two factors: high-vacuity samples are prioritized early on to ensure coverage of underrepresented cases, while high-dissonance samples are prioritized later to refine decision boundaries. This adaptive approach balances exploration and refinement for efficient labeling.
- Vacuity (Vac(x))
- Vacuity measures the lack of evidence for a specific sample, calculated by dividing a constant by the sum of similarity scores across classes. High vacuity flags rare or unseen phenotypes where the model lacks sufficient total evidence to make a confident judgment.
Terminology used across episodes
This episode discusses
- Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning · Paper Radio
- Lung and Colon Cancer Histopathological Image Dataset (LC25000)
- Enabling Calibration In The Zero-Shot Inference of Large Vision-Language Models
- A Closer Look at the Explainability of Contrastive Language-Image Pre-training
- A Dynamic Temporal Self-attention Graph Convolutional Network for Traffic Prediction
The paper
Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning · Read on arXiv
School of Electronic Technology and Engineering, Xiamen University · School of Informatics, Xiamen University · School of Engineering, Case Western Reserve University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning".
Tom: The gist The Similarity-as-Evidence (SaE) framework calibrates text–image similarities by introducing a Similarity Evidence Head (SEH),
Jane: First, who's behind it and why it matters.
Title and authors: Tom: Let’s go over the paper itself, Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning. The authors are Zhuofan Xie, Zishan Lin, Jinliang Lin, Jie Qi, Shaohua Hong, and Shuo Li.
Jane: It's a team from Xiamen University and Case Western Reserve University who are tackling this overconfidence issue head-on in the context of medical active learning.
Lu: The core idea they present is that instead of just using the similarity vector directly, you should reinterpret that similarity as evidence and parameterize a Dirichlet distribution over labels.
Meng: So instead of a single score, we get parameters for how much belief the model has in different possible outcomes based on that image. That’s a big shift from what we see in standard setups.
Tom: It changes how we think about the similarity vector entirely, moving it from just a certainty measure to something that quantifies actual evidence strength.
Jane: It suggests that this approach can lead to selection rationales for when we ask for more labels that are much clearer clinically.
The paper's summary: Tom: To get into the summary, the paper explains how they use this Similarity Evidence Head to decompose the evidence into two things: vacuity and dissonance.
Jane: Vacuity is basically a measure of knowledge gaps, flagging those rare phenotypes where there isn't enough total evidence for the model to make a solid judgment on it yet.
Lu: And dissonance measures decision conflicts, which means when the model is pulling conflicting ideas across different classes when looking at that image.
Meng: That’s useful because it separates whether we need more data because we don't know enough about something, or if the model is just confused between two possibilities.
Tom: The paper sets up a loss function called LSEH which balances two things: it forces the evidence strength to match how hard the ground truth is, while also keeping that strength consistent with what the VLM already knows.
Jane: So they are training this head not just to be accurate, but also to be calibrated in a way that makes sense for clinical decision-making.
The paper's improvements: Tom: The real improvements they suggest are in how we structure the acquisition strategy based on those vacuity and dissonance factors. They propose a dual-factor acquisition strategy.
Lu: In the early rounds, you prioritize samples with high vacuity to make sure you cover all the different cases and phenotypes in our dataset.
Jane: That way, we ensure good coverage of underrepresented things before we get too focused on refining the boundaries.
Meng: And later in the rounds, when we move into refinement, you switch to prioritizing samples with high dissonance so you can focus on those hard-to-distinguish cases that are actually confusing the model.
Tom: That adaptive scheduling, weighting vacuity early and dissonance later based on round number, is what they claim outperforms static acquisition strategies.
Jane: They also found that setting a context vector length of sixteen yields the best accuracy across both datasets, which is a practical tuning tip for implementation.
Conclusion: Tom: So, to wrap up this discussion on Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning. The main point is moving away from raw similarity scores to evidence quantified by the SEH.
Jane: It gives us a way to select samples that is not just based on a high score, but on a calibrated measure of how much knowledge we actually gain from labeling that specific sample.
Lu: This framework helps bridge the gap between general VLM alignment and the specific semantics of medical imaging by making the evidence more structured.
Meng: For practical deployment, it means our annotation budget gets spent smarter because we’re targeting both ignorance and confusion in a predictable way rather than just chasing high scores blindly.
Lalam: I think this approach is really important for improving how we build systems that interact with medical data, as it gives us a more robust way to handle the inherent uncertainty in diagnostic tasks.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck