Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning

summary

Video file (mp4)

The gist

The gist The Similarity-as-Evidence (SaE) framework calibrates text–image similarities by introducing a Similarity Evidence Head (SEH), which reinterprets the similarity vector as evidence and

In short

The Similarity-as-Evidence (SaE) framework recalibrates overconfident Vision-Language Models for medical active learning by treating similarity scores as evidence rather than deterministic predictions. It introduces a Similarity Evidence Head to quantify model certainty, decomposes uncertainty into knowledge gaps (vacuity) and decision conflicts (dissonance), and uses a dual-factor acquisition strategy to select the most informative samples, leading to improved accuracy and calibration.

Key concepts

Similarity Evidence Head (SEH)
This component maps raw text-image similarity scores into Dirichlet evidence parameters. It is trained with a loss function that balances classification performance against the model's intrinsic certainty, allowing the framework to quantify how much evidence the VLM has for each prediction.
Uncertainty Decomposition
The framework breaks down uncertainty into two parts: vacuity, which represents knowledge gaps or missing information about rare phenotypes, and dissonance, which measures conflicts between competing class hypotheses. This decomposition provides a more nuanced understanding of why a model is uncertain.
Dual-Factor Acquisition Strategy
This strategy selects samples based on two factors: high-vacuity samples are prioritized early on to ensure coverage of underrepresented cases, while high-dissonance samples are prioritized later to refine decision boundaries. This adaptive approach balances exploration and refinement for efficient labeling.
Vacuity (Vac(x))
Vacuity measures the lack of evidence for a specific sample, calculated by dividing a constant by the sum of similarity scores across classes. High vacuity flags rare or unseen phenotypes where the model lacks sufficient total evidence to make a confident judgment.

Terminology used across episodes

This episode discusses

The paper

Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning · Read on arXiv

School of Electronic Technology and Engineering, Xiamen University · School of Informatics, Xiamen University · School of Engineering, Case Western Reserve University

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning".

Tom: The gist The Similarity-as-Evidence (SaE) framework calibrates text–image similarities by introducing a Similarity Evidence Head (SEH),

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Let’s go over the paper itself, Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning. The authors are Zhuofan Xie, Zishan Lin, Jinliang Lin, Jie Qi, Shaohua Hong, and Shuo Li.

Jane: It's a team from Xiamen University and Case Western Reserve University who are tackling this overconfidence issue head-on in the context of medical active learning.

Lu: The core idea they present is that instead of just using the similarity vector directly, you should reinterpret that similarity as evidence and parameterize a Dirichlet distribution over labels.

Meng: So instead of a single score, we get parameters for how much belief the model has in different possible outcomes based on that image. That’s a big shift from what we see in standard setups.

Tom: It changes how we think about the similarity vector entirely, moving it from just a certainty measure to something that quantifies actual evidence strength.

Jane: It suggests that this approach can lead to selection rationales for when we ask for more labels that are much clearer clinically.

The paper's summary: Tom: To get into the summary, the paper explains how they use this Similarity Evidence Head to decompose the evidence into two things: vacuity and dissonance.

Jane: Vacuity is basically a measure of knowledge gaps, flagging those rare phenotypes where there isn't enough total evidence for the model to make a solid judgment on it yet.

Lu: And dissonance measures decision conflicts, which means when the model is pulling conflicting ideas across different classes when looking at that image.

Meng: That’s useful because it separates whether we need more data because we don't know enough about something, or if the model is just confused between two possibilities.

Tom: The paper sets up a loss function called LSEH which balances two things: it forces the evidence strength to match how hard the ground truth is, while also keeping that strength consistent with what the VLM already knows.

Jane: So they are training this head not just to be accurate, but also to be calibrated in a way that makes sense for clinical decision-making.

The paper's improvements: Tom: The real improvements they suggest are in how we structure the acquisition strategy based on those vacuity and dissonance factors. They propose a dual-factor acquisition strategy.

Lu: In the early rounds, you prioritize samples with high vacuity to make sure you cover all the different cases and phenotypes in our dataset.

Jane: That way, we ensure good coverage of underrepresented things before we get too focused on refining the boundaries.

Meng: And later in the rounds, when we move into refinement, you switch to prioritizing samples with high dissonance so you can focus on those hard-to-distinguish cases that are actually confusing the model.

Tom: That adaptive scheduling, weighting vacuity early and dissonance later based on round number, is what they claim outperforms static acquisition strategies.

Jane: They also found that setting a context vector length of sixteen yields the best accuracy across both datasets, which is a practical tuning tip for implementation.

Conclusion: Tom: So, to wrap up this discussion on Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning. The main point is moving away from raw similarity scores to evidence quantified by the SEH.

Jane: It gives us a way to select samples that is not just based on a high score, but on a calibrated measure of how much knowledge we actually gain from labeling that specific sample.

Lu: This framework helps bridge the gap between general VLM alignment and the specific semantics of medical imaging by making the evidence more structured.

Meng: For practical deployment, it means our annotation budget gets spent smarter because we’re targeting both ignorance and confusion in a predictable way rather than just chasing high scores blindly.

Lalam: I think this approach is really important for improving how we build systems that interact with medical data, as it gives us a more robust way to handle the inherent uncertainty in diagnostic tasks.

More episodes

← Home