Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning".
Jane: Pathology foundation models (PFMs) offer generalizable representations for whole-slide image (WSI) analysis,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: Welcome back to the show! Today we're looking at some really interesting work coming out of arXiv that’s touching on how we can make sure AI models, especially in complex areas like medical imaging, don't just give answers without telling us how sure they are. We’re talking about the paper "Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning."
Jane: That sounds super practical, Tom; it seems like they are trying to solve a real headache for anyone using these huge foundation models for things like whole-slide image analysis.
Lu: Yeah, the abstract sets up a clear problem: these foundation models give us general representations for whole-slide images, but clinically they struggle because their predictions lack reliable confidence estimates and there isn't one single best model across different tasks <ref:2606.30020#pg1>. That lack of confidence is a huge hurdle when you're dealing with medical decisions.
Meng: From an engineering viewpoint, that uncertainty issue is critical because in a real clinical setting, we absolutely need to know when a model is unsure before we trust its output for diagnosis <ref:2606.30020#pg1>. We can't just let it give us a result without a confidence score attached.
Lalam: I think this paper is really important because it’s about building trust into the AI systems used for diagnosis, which is something we need to focus on as we integrate these tools more deeply into our lives <ref:2606.30020#pg1>. It moves the conversation from just getting an output to understanding the reliability behind that output.
Tom: So they propose DICE, which is this plug-and-play framework that uses an ensemble of K frozen PFMs and models their disagreement as a proxy for uncertainty estimation <ref:2606.30020#pg1>. It seems the central claim here is that the disagreement among these different PFMs acts as a meaningful signal for estimating the actual uncertainty in what they predict.
Jane: That makes sense, Tom; so they align these experts using deep mutual learning to make sure their disagreements actually reflect real uncertainty, and they theoretically show that this alignment process provides an upper bound on the model uncertainty <ref:2606.30020#pg1>.
Lu: That deep mutual learning objective, which involves minimizing a loss term related to KL divergence between expert posteriors, is what establishes that theoretical guarantee about bounding the model uncertainty <ref:2606.30020#pg1>. It’s a solid mathematical underpinning for why this approach should work.
Meng: I wonder about the practical side of running K frozen models and managing all that alignment during training; how feasible is it to do this without making the whole process too slow for real-world use cases <ref:2606.30020#pg1>? We need something deployable, not just a complex research setup.
Lalam: It’s about creating a system where the ensemble itself provides that uncertainty quantification, which I think is a really cool way to make the AI more transparent for users <ref:2606.30020#pg1>. Transparency is key when we're talking about tools that help guide clinical decisions.
Paper summary: Tom: And they don't stop there; they also show that this expert consensus can actually help with localization, which is another quite interesting finding in the paper <ref:2606.30020#pg2>. This goes beyond just knowing *what* something is; it helps them pinpoint *where* it is on the slide.
Jane: I remember reading about how they use an attention map consensus to help pinpoint abnormalities at the patch level without needing any extra supervision for that localization task <ref:2606.30020#pg2>. That ability to localize things automatically is a big deal for diagnostics.
Lu: That localization aspect is significant because it suggests that the agreement among experts isn't just about the final classification label; it translates into a sharp signal for where the pathology actually resides, even without explicit labels for that localization task <ref:2606.30020#pg2>.
Tom: So, to wrap up this summary of "Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning," we see they are shifting away from relying on just one model to provide all the confidence information <ref:2606.30020#pg1>. The authors suggest that by treating disagreements between PFMs as an informative signal, they can translate these models into systems that are aware of their own uncertainty.
Jane: It seems the core message is that by leveraging these diverse AI models and modeling their internal conflicts, we can build a more robust framework for understanding when an AI might be making a guess <ref:2606.30020#pg1>.
Meng: From an engineering perspective, the practical implication here is improving the reliability of whole-slide image analysis by giving us more reliable confidence scores, which directly helps in informed model selection for clinical use <ref:2606.30020#pg1>. That’s something we can measure and deploy.
Lalam: This work really shows how we can leverage diverse AI models to create a more robust and trustworthy diagnostic tool, moving us closer to systems that are honest about their limitations <ref:2606.30020#pg1>.
Tom: Moving on to the conclusion of this discussion regarding "Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning," the title itself really captures the essence of what they achieved <ref:2606.30020#pg1>. The authors framed their work around using deep mutual learning for alignment and focusing squarely on how to estimate uncertainty in pathology models, which is a smart way to frame this complex research.
Jane: It’s about taking these very large foundation models and figuring out exactly how to quantify their uncertainty so that we can actually trust them more in a clinical environment where mistakes have serious consequences <ref:2606.30020#pg1>.
Lu: I think the authors chose that title because it highlights the two main pillars of their work: using deep mutual learning for alignment and focusing on uncertainty estimation, which is a really smart way to frame this research <ref:2606.30020#pg1>. It sets up a very clear roadmap for what they accomplished.
Meng: From my side, I'm thinking about how that title speaks to moving beyond just getting an answer from a model and actually understanding the confidence behind that answer, which is exactly what we need for real-world deployment <ref:2606.30020#pg1>. We need to know when the uncertainty is high enough to flag a case for human review.
Paper summary: Lalam: I feel that title perfectly reflects how they are taking general knowledge from these models and making it trustworthy by adding this layer of uncertainty quantification, which could really help in building better cultural tools based on medical images <ref:2606.30020#pg1>. It’s about making the AI honest.
Tom: Right, and the authors are clearly smart; they've managed to tackle a problem where relying on a single model falls short, showing how ensembling and mutual learning can solve that issue <ref:2606.30020#pg1>. The entire framework is quite clever in how it handles the complexity of multiple experts working together.
Jane: They’ve shown that this approach isn't just a clever trick but is grounded in deep alignment techniques, which makes the resulting uncertainty estimates much more reliable than just picking one model at random <ref:2606.30020#pg1>. The rigor behind the method is what gives us confidence in the results.
Lu: The deep mutual learning objective they use to align the experts really establishes a solid theoretical foundation for how we can contract that uncertainty bound, which is pretty profound research <ref:2606.30020#pg1>. That theoretical proof is what makes this work more than just empirical observation.
Meng: I’m just thinking about the practical implications of that contraction; if the math holds up, it means we have a predictable way to manage model risk when we deploy these systems in a live clinical setting <ref:2606.30020#pg1>. That predictability is what engineers look for most.
Lalam: For me, this implies that as AI becomes more integrated into our daily lives through things like medical diagnostics, having a formal way to say "I'm uncertain about this" could really improve how we interact with and trust those systems in general <ref:2606.30020#pg1>.
Tom: Exactly. So, the main point we’re taking away from "Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning" is that ensembling and mutual learning provide a structured way to quantify uncertainty that goes beyond what any single model can offer <ref:2606.30020#pg1>.
Jane: It’s about taking those powerful foundation models and figuring out how to measure their internal disagreement so we can build systems that are aware of when they might be guessing <ref:2606.30020#pg1>.
Lu: The implication here for the broader field is that we move toward task-agnostic uncertainty quantification, which was something that wasn't really there before because previous methods were often too specific to one particular task or model architecture <ref:2606.30020#pg1>. That’s a big conceptual step forward.
Meng: From an engineering view, it suggests a way to deploy multiple models and use their internal dynamics to gauge reliability dynamically during inference, which is something we can actually start prototyping now <ref:2606.30020#pg1>.
Lalam: If this framework works well, it means we can develop AI tools that are not just accurate but also honest about when they might be guessing, which is a huge step toward real clinical integration and better user interaction <ref:2606.30020#pg1>.
Conclusion: Tom: So we've seen how DICE uses frozen PFMs to estimate uncertainty by modeling their disagreements, and now we're coming to the finish line of this discussion about its title and authors <ref:2606.30020#pg1>.
Jane: Exactly, Tom; the paper is titled "Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning," and it really captures how they take these big foundation models and build a system around them to figure out when they're unsure <ref:2606.30020#pg1>.
Lu: I think the authors chose that title because it highlights the two main pillars of their work: using deep mutual learning for alignment and focusing on uncertainty estimation, which is a really smart way to frame the research <ref:2606.30020#pg1>.
Meng: From my side, I'm thinking about how this title speaks to moving beyond just getting an answer from a model and actually understanding the confidence behind that answer, which is what we need for real-world deployment <ref:2606.30020#pg1>.
Lalam: I feel that title perfectly reflects how they're taking general knowledge and making it trustworthy by adding this layer of uncertainty quantification, which could really help in building better cultural tools based on medical images <ref:2606.30020#pg1>.
Tom: Right, and the authors are clearly smart; they've managed to tackle a problem where single models fall short, showing how ensembling and mutual learning can solve it <ref:2606.30020#pg1>.
Jane: They’ve shown that this approach isn't just a clever trick but is grounded in deep alignment techniques, which makes the uncertainty estimates much more reliable than just guessing <ref:2606.30020#pg1>.
Lu: The deep mutual learning objective they use to align the experts really establishes a solid theoretical foundation for how we can contract that uncertainty bound, which is pretty profound research <ref:2606.30020#pg1>.
Meng: I’m just thinking about the practical implications of that contraction; if the math holds up, it means we have a predictable way to manage model risk when deploying these systems <ref:2606.30020#pg1>.
Lalam: For me, this implies that as AI becomes more integrated into our daily lives through things like medical diagnostics, having a formal way to say "I'm uncertain about this" could really improve how we interact with and trust those systems <ref:2606.30020#pg1>.
Tom: So the core message here is that by using an ensemble and deep learning to model disagreements, DICE gives us a structured way to quantify uncertainty that goes way beyond what any single model can offer <ref:2606.30020#pg1>.
Jane: It’s about taking those powerful foundation models and figuring out exactly how to measure their internal conflicts so we can build systems that are aware of when they might be guessing <ref:2606.30020#pg1>.
Lu: The implication for the broader field is that we move toward task-agnostic uncertainty quantification, which was something that wasn't really there before because previous methods were often too specific to one particular task or model architecture <ref:2606.30020#pg1>.
Meng: From an engineering view, it suggests a way to deploy multiple models and use their internal dynamics to gauge reliability dynamically during inference, which is something we can actually start prototyping now <ref:2606.30020#pg1>.
Lalam: If this framework works well, it means we can develop AI tools that are not just accurate but also honest about when they might be guessing, which is a huge step toward real clinical integration <ref:2606.30020#pg1>.
Tom: This work really shows how ensembling and mutual learning provide a structured way to quantify uncertainty that goes beyond what any single model can offer <ref:2606.30020#pg1>.
Jane: It’s about taking those powerful foundation models and figuring out exactly how to measure their internal conflicts so we can build systems that are aware of when they might be guessing <ref:2606.30020#pg1>.
Lu: The implication here for the broader field is that we move toward task-agnostic uncertainty quantification, which was something that wasn't really there before because previous methods were often too specific to one particular task or model architecture <ref:2606.30020#pg1>.
Gbègninougbo Aurel Davy Tchokponhoue, Sevda Ögüt, Ali Idri, Dorina Thanou, Pascal Frossard
UM6P, Ben Guerir, Morocco · EPFL, Lausanne, Switzerland
cs.CV
Submitted: 2026-06-29
Updated: 2026-10-02
Code: https://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution
Importance score: 81/100
The gist: Pathology foundation models (PFMs) offer generalizable representations for whole-slide image (WSI) analysis, yet their clinical adoption remains limited because their predictions lack reliable
Key concepts
- Deep Mutual Learning (DML)
- This is a training objective used to align different PFM experts. It forces each expert's prediction distribution to match those of its peers while still staying anchored to the correct ground-truth label. This ensures that the ensemble members learn similar, coherent ways of making predictions.
- Gramian Measure
- This loss function is used during training to ensure that the internal representations (embeddings) of different PFM experts are spread out across a wide range of possibilities. It encourages the models to cover a low-volume subspace, which helps in detecting genuine data uncertainty rather than just representation mismatch.
- Epistemic Uncertainty
- This type of uncertainty measures what the ensemble doesn't know because the experts disagree on how to interpret the input. It is formally calculated as the divergence between expert posteriors. High epistemic uncertainty signals a sample where the models are genuinely confused, making it a reliable indicator for decision-making.
- Patch-level Consensus
- This metric assesses how well different experts agree on where an abnormality is located at the small patch level. It is calculated by averaging the attention coefficients from all experts across that patch. A sharp signal here indicates a highly localized, diagnostically relevant region.
Terminology
Summary
Pathology foundation models (PFMs) offer generalizable representations for whole-slide image (WSI) analysis, yet their clinical adoption remains limited because their predictions lack reliable confidence estimates and no single PFM is universally best across tasks. To address this critical gap, DICE proposes a plug-and-play framework that ensembles K frozen PFMs and models their disagreement as a proxy for uncertainty estimation, thereby translating PFMs into uncertainty-aware decision-support systems.
The gist
DICE is a plug-and-play framework that ensembles K frozen PFMs and models their disagreement as a proxy for uncertainty estimation, aligning the ensemble members via deep mutual learning to provide task-agnostic uncertainty quantification in computational pathology.
How it works
The core intuition behind DICE is that disagreement among heterogeneous PFMs is an informative signal for uncertainty estimation.
The framework treats each PFM and its corresponding classification head as an expert, and these experts are aligned through two primary objectives during training:
-
Prediction-space alignment via deep mutual learning (DML). Each expert minimizes a loss function that includes the standard cross-entropy loss on the ground truth label, alongside a mutual distillation term:
L(k)DML = 1/(K − 1) X l ≠ k KL p(l) / p(k)
. This encourages expert k toalign its posterior with those of its peers while remaining anchored to the ground-truth label through the supervised loss.
-
Representation-space alignment via Gramian measure. To ensure that disagreement reflects genuine data uncertainty rather than representation mismatch, a Gramian loss is added:
LGram = r det R˜⊤R˜,
which encourages expert slide embeddings tospan a low-volume subspace.
Uncertainty Representation
At inference, the ensemble prediction is obtained by averaging the predictive distributions of the K experts: p¯(B) = 1/K Σ p(k)(B). This prediction is then decomposed into its components: u pred = H(p¯z) Total Uncertainty,
u alea = 1/K X K k=1 H[p(k)(B)z] Data Uncertainty,
and u epi = u pred − u alea model Uncertainty.
The epistemic uncertainty, which captures the residual disagreement across experts, is formally expressed as the multi-way Jensen–Shannon divergence among expert posteriors: u epi = 1/K X K k=1 KL p(k) p¯ = JSD p(1),..., p(K).
Proposition 2 establishes a formal upper bound on this epistemic uncertainty: u epi ≤ (K − 1)/K squared LDML.
Empirical Evaluation and Localization
DICE was evaluated on three challenging WSI benchmarks: PANDA, CAMELYON16, and CAMELYON17. The framework demonstrated that experts’ residual disagreement identifies uncertain samples more reliably than standard baselines, in both in- and out-of-distribution settings.
Furthermore, the ensemble exploits expert consensus to improve localization: expert consensus localizes abnormalities at the patch level without any explicit supervision.
Patch-level lesion localization was assessed by defining patch-level consensus as the empirical mean of expert attention coefficients (c patch i = 1/K Σ a(k) i). DICE showed that this signal is a sharp localization signal,
concentrating attention inside annotated tumor regions.
Performance and Generalization
Empirically, DICE outperformed baselines across classification, calibration, and localization metrics. On PANDA, it delivered the largest gains in F1 score over the best single PFM. In terms of uncertainty estimation reliability, DICE showed superior performance in selective-prediction experiments: deferring high-entropy slides monotonically lowers the retained error rate.
Moreover, DICE’s uncertainty signal generalized across data splits and cohorts; under a cohort shift on CAMELYON17, DICE (w/o reg) achieves the strongest separation with ∆ = 0.40,
indicating that expert disagreement remains informative even when models are trained on different datasets. The framework successfully matches or exceeds the strongest single PFM baselines across a range of downstream tasks and metrics.
Contributions
The main contributions include:
-
Proposing DICE, the
first plug-and-play framework that fuses PFM ensembles for uncertainty-aware decision support in histopathology.
-
Providing a theoretical foundation by proving that
minimizing the training-time DML alignment objective contracts an upper bound on the model uncertainty.
-
Demonstrating that DICE not only
accurately highlights unreliable samples, but also improves classification, calibration, and localization performance compared to SOTA models.
-
Showing that attention consensus is a
sharp localization signal,
enabling faster identification of diagnostically relevant tissue regions without explicit supervision.
Improvements for AI systems
Based on the scientific paper Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning
(DICE), here are specific improvements to AI systems that can be made, focusing on integrating DICE into existing Pathology Foundation Model (PFM) pipelines:
-
Improve Clinical Trust and Safety by Implementing Uncertainty-Aware Decision Support:
-
Enhance Model Selection for Deployment by Quantifying PFM Reliability:
-
Optimize Lesion Detection Accuracy through Consensus-Based Localization:
-
Develop Robust Out-of-Distribution (OOD) Detection Capabilities in Unseen Pathology Data:
Here is a detailed breakdown of what the improved AI system can do, derived directly from the paper's contributions:
-
The improved system can provide a reliable, principled measure of confidence for every whole-slide image (WSI) analysis output. By leveraging the ensemble disagreement mechanism (specifically residual disagreement bounded by Deep Mutual Learning loss), it can flag
failure-prone
slides—cases where the model is likely uncertain or making an error—allowing clinicians to defer these ambiguous cases for manual review, thereby significantly enhancing patient safety in clinical settings. -
The system can move beyond relying on a single PFM by providing a framework for selecting the most appropriate PFM configuration for a specific task (e.g., grading vs. metastasis detection). Because DICE explicitly models the disagreement and performance across different architectures and pretraining objectives, it allows researchers to select the ensemble that is theoretically best suited for the current clinical context, leading to more robust and reliable downstream predictions than relying on a single
best
model. -
The system can improve the localization of tissue abnormalities (lesions) by exploiting expert consensus. Instead of relying solely on the attention map of one PFM, DICE uses the patch-level consensus (the empirical mean of all K expert attention coefficients). This results in sharper, more accurate localization signals that concentrate attention precisely within annotated tumor regions, which is critical for identifying small lesions that are often missed by single models.
-
The system can effectively detect and manage uncertainty when encountering data outside its training distribution (OOD settings). By treating the ensemble's disagreement as a proxy for uncertainty, DICE is shown to be reliable in flagging cases that deviate from the expected data distribution, ensuring that the system behaves predictably and safely even when presented with novel or unusual pathology slides.
Sources
- Ensemble learning of pathology foundation models for precision oncology
- Hibou: A Family of Foundational Vision Transformers for Pathology
- GrapHist: Graph Self-Supervised Learning for Histopathology
- PRISM: A Multi-Modal Generative Foundation Model for Slide-Level Histopathology
- Fusion of Multi-scale Heterogeneous Pathology Foundation Models for Whole Slide Image Analysis
- Accelerating Data Processing and Benchmarking of AI Models for Pathology
- Virchow2: Scaling Self-Supervised Mixed Magnification Models in Pathology
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models