Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning".
Tom: The gist The Similarity-as-Evidence (SaE) framework calibrates text–image similarities by introducing a Similarity Evidence Head (SEH),
Jane: First, who's behind it and why it matters.
Title and authors: Tom: Let’s go over the paper itself, Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning. The authors are Zhuofan Xie, Zishan Lin, Jinliang Lin, Jie Qi, Shaohua Hong, and Shuo Li.
Jane: It's a team from Xiamen University and Case Western Reserve University who are tackling this overconfidence issue head-on in the context of medical active learning.
Lu: The core idea they present is that instead of just using the similarity vector directly, you should reinterpret that similarity as evidence and parameterize a Dirichlet distribution over labels.
Meng: So instead of a single score, we get parameters for how much belief the model has in different possible outcomes based on that image. That’s a big shift from what we see in standard setups.
Tom: It changes how we think about the similarity vector entirely, moving it from just a certainty measure to something that quantifies actual evidence strength.
Jane: It suggests that this approach can lead to selection rationales for when we ask for more labels that are much clearer clinically.
The paper's summary: Tom: To get into the summary, the paper explains how they use this Similarity Evidence Head to decompose the evidence into two things: vacuity and dissonance.
Jane: Vacuity is basically a measure of knowledge gaps, flagging those rare phenotypes where there isn't enough total evidence for the model to make a solid judgment on it yet.
Lu: And dissonance measures decision conflicts, which means when the model is pulling conflicting ideas across different classes when looking at that image.
Meng: That’s useful because it separates whether we need more data because we don't know enough about something, or if the model is just confused between two possibilities.
Tom: The paper sets up a loss function called LSEH which balances two things: it forces the evidence strength to match how hard the ground truth is, while also keeping that strength consistent with what the VLM already knows.
Jane: So they are training this head not just to be accurate, but also to be calibrated in a way that makes sense for clinical decision-making.
The paper's improvements: Tom: The real improvements they suggest are in how we structure the acquisition strategy based on those vacuity and dissonance factors. They propose a dual-factor acquisition strategy.
Lu: In the early rounds, you prioritize samples with high vacuity to make sure you cover all the different cases and phenotypes in our dataset.
Jane: That way, we ensure good coverage of underrepresented things before we get too focused on refining the boundaries.
Meng: And later in the rounds, when we move into refinement, you switch to prioritizing samples with high dissonance so you can focus on those hard-to-distinguish cases that are actually confusing the model.
Tom: That adaptive scheduling, weighting vacuity early and dissonance later based on round number, is what they claim outperforms static acquisition strategies.
Jane: They also found that setting a context vector length of sixteen yields the best accuracy across both datasets, which is a practical tuning tip for implementation.
Conclusion: Tom: So, to wrap up this discussion on Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning. The main point is moving away from raw similarity scores to evidence quantified by the SEH.
Jane: It gives us a way to select samples that is not just based on a high score, but on a calibrated measure of how much knowledge we actually gain from labeling that specific sample.
Lu: This framework helps bridge the gap between general VLM alignment and the specific semantics of medical imaging by making the evidence more structured.
Meng: For practical deployment, it means our annotation budget gets spent smarter because we’re targeting both ignorance and confusion in a predictable way rather than just chasing high scores blindly.
Lalam: I think this approach is really important for improving how we build systems that interact with medical data, as it gives us a more robust way to handle the inherent uncertainty in diagnostic tasks.
School of Electronic Technology and Engineering, Xiamen University · School of Informatics, Xiamen University · School of Engineering, Case Western Reserve University
cs.CV
Submitted: 2026-02-21
Updated: 2026-10-08
Comments: Accepted to CVPR 2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 87/100
The gist: The gist The Similarity-as-Evidence (SaE) framework calibrates text–image similarities by introducing a Similarity Evidence Head (SEH), which reinterprets the similarity vector as evidence and
Key concepts
- Similarity Evidence Head (SEH)
- This component maps raw text-image similarity scores into Dirichlet evidence parameters. It is trained with a loss function that balances classification performance against the model's intrinsic certainty, allowing the framework to quantify how much evidence the VLM has for each prediction.
- Uncertainty Decomposition
- The framework breaks down uncertainty into two parts: vacuity, which represents knowledge gaps or missing information about rare phenotypes, and dissonance, which measures conflicts between competing class hypotheses. This decomposition provides a more nuanced understanding of why a model is uncertain.
- Dual-Factor Acquisition Strategy
- This strategy selects samples based on two factors: high-vacuity samples are prioritized early on to ensure coverage of underrepresented cases, while high-dissonance samples are prioritized later to refine decision boundaries. This adaptive approach balances exploration and refinement for efficient labeling.
- Vacuity (Vac(x))
- Vacuity measures the lack of evidence for a specific sample, calculated by dividing a constant by the sum of similarity scores across classes. High vacuity flags rare or unseen phenotypes where the model lacks sufficient total evidence to make a confident judgment.
Terminology
Summary
The gist The Similarity-as-Evidence (SaE) framework calibrates text–image similarities by introducing a Similarity Evidence Head (SEH), which reinterprets the similarity vector as evidence and parameterizes a Dirichlet distribution over labels, thereby mitigating overconfidence caused by rigid softmax normalization
Problem Addressed
Active Learning (AL) suffers from a cold-start problem when labeled data are scarce, which is exacerbated by Vision-Language Models (VLMs) being inherently overconfident due to their temperature-scaled softmax outputs treating text–image similarities as deterministic scores while ignoring inherent uncertainty This overconfidence misleads sample selection, wasting annotation budgets on uninformative cases Existing calibration methods provide only global adjustments and do not explain why a prediction is uncertain Furthermore, most AL strategies rely on scalar uncertainty scores such as predictive entropy or margin, which do not reveal whether uncertainty arises from missing knowledge or from conflicting hypotheses
Similarity-as-Evidence (SaE) Framework
The SaE framework recalibrates VLM outputs into interpretable evidential uncertainty for medical AL by introducing three synergistic components
-
Similarity Evidence Head (SEH): This component maps raw text–image similarities to Dirichlet evidence parameters, quantifying how much evidence the VLM has for each prediction It is trained with a dual-objective loss that balances classification performance and uncertainty calibration
-
Uncertainty Decomposition: The evidence is decomposed into vacuity (knowledge gaps) and dissonance (decision conflicts) via Subjective Logic (SL)
-
Dual-Factor Acquisition Strategy: This strategy prioritizes high-vacuity samples in early rounds to ensure coverage, while high-dissonance samples are prioritized later to refine boundaries
Key Components of SaE
The SEH estimates a scalar strictly positive evidence strength λ by fusing image features and similarity scores through a dual-branch MLP architecture The SEH Loss Function, LSEH, is a dual-objective loss function that balances empirical classification difficulty with the VLM’s intrinsic certainty The first term, Ldiff, forces the evidence strength to reflect groundtruth difficulty (1/λ ↑ for hard samples with high lcls), while the second term, Lent, ensures consistency with the VLM’s prior knowledge (low entropy ⇒ high λ) The acquisition score for each sample x is a dynamic weighted combination of its normalized uncertainty factors: Score t(x) = w v(t) · g(Vac(x)) + w d(t) · g(Dis(x))
Dual-Factor Acquisition Strategy
The dual-factor strategy decomposes the evidential uncertainty into two clinically meaningful factors under SL
** Vacuity (lack of evidence): defined as Vac(x) = K / Σαk(x), which flags rare or currently unseen phenotypes where the model lacks sufficient total evidence to make a judgment The acquisition score is weighted by w v(t) = 1 - (t-1)/(T-1) in early rounds to prioritize high-vacuity samples This ensures coverage of under-represented phenotypes The dissonance factor, Dis(x), measures the conflict among competing classes using a complex formula involving belief masses bk(x) The acquisition score is weighted by w d(t) = (t-1)/(T-1) in later rounds to focus on high-dissonance, hard-to-distinguish cases This adaptive strategy aligns sample selection with clinical reasoning and provides clinically interpretable rationales for annotation requests The dynamic schedule is shown to outperform all static variants, optimally bridging the needs of exploration and refinement The SaE framework attains state-of-the-art macro-averaged accuracy of 82.57% on ten public medical imaging datasets with a 20% label budget On the representative BTMRI dataset, SaE also achieves superior calibration, with a negative loglikelihood (NLL) of 0.425 The framework's performance is robust across a wide range of the loss weight β, with the optimal trade-off consistently observed at β = 0.5 A moderate length (M = 16) for context vectors yields the best accuracy by balancing semantic capacity and overfitting This suggests that setting M = 16 consistently yields the optimal performance across both datasets The dynamic schedule outperforms all static variants, while the dissonance-only strategy yields the poorest results (89.12%) The SaE framework successfully localizes clinical features, generating highly focused attention maps that align closely with the ground-truth lesion contours This visual evidence supports the claim that the evidential calibration mechanism guides the model to leverage correct semantic features The overall process is summarized in Algorithm 1 The results show that SaE consistently outperforms existing VLM-based AL methods in both accuracy and calibration Overall, SaE advances interpretable, label-efficient AL and strengthens the reliability of VLM-driven pipelines for clinical deployment The research was funded by the Key Program of Marine Economy Development Special Foundation of Department of Natural Resources of Guangdong Province (GDNRC[2023]24) The paper is arXiv: 2602.18867v2 [cs.
Improvements for AI systems
-
A Similiarity-as-Evidence (SaE) framework calibrates text–image similarities by introducing a Similarity Evidence Head (SEH), which
reinterprets the similarity vector as evidence and parameterizes a Dirichlet distribution over labels,
enablingclinically interpretable selection rationales.
This allows the AI to select samples based on decomposed uncertainty intovacuity (knowledge gaps)
anddissonance (decision conflicts).
-
The dual-factor acquisition strategy enables adaptive sample selection: high-vacuity samples are prioritized in early rounds
to ensure coverage,
while high-dissonance samples are prioritized laterto refine boundaries,
providing a mechanism that aligns the acquisition function with clinical reasoning. -
The SEH loss function, which balances empirical difficulty and VLM confidence, is trained using a dual-objective loss:
Ldiff forces the evidence strength to reflect groundtruth difficulty
whileLent ensures consistency with the VLM’s prior knowledge,
resulting in calibrated evidence strength λ that mitigates overconfidence. -
The system can generate visually interpretable uncertainty maps using Grad-CAM, where SaE produces
highly focused and accurate attention maps that align closely with the ground-truth lesion contours,
confirming that itsevidence-calibrated strategy successfully localizes clinical features.
-
The AI system can be adapted to specific medical domains by augmenting class prompts with specialized knowledge retrieved from PubMed, creating a
semantically rich similarity vector s
which serves as the input to the evidence model, bridging the gap between general VLM knowledge andthe specific semantics of medical imaging.
Sources
- Lung and Colon Cancer Histopathological Image Dataset (LC25000)
- Enabling Calibration In The Zero-Shot Inference of Large Vision-Language Models
- A Closer Look at the Explainability of Contrastive Language-Image Pre-training
- A Dynamic Temporal Self-attention Graph Convolutional Network for Traffic Prediction
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models