Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning

summary

Video file (mp4)

The gist

Pathology foundation models (PFMs) offer generalizable representations for whole-slide image (WSI) analysis, yet their clinical adoption remains limited because their predictions lack reliable

In short

DICE is a plug-and-play framework that combines multiple frozen pathology foundation models (PFMs) into an ensemble to estimate uncertainty. It aligns these experts through deep mutual learning and representation alignment, using their disagreement as a proxy for uncertainty. This allows the system to provide reliable confidence estimates for whole-slide image analysis.

Key concepts

Deep Mutual Learning (DML)
This is a training objective used to align different PFM experts. It forces each expert's prediction distribution to match those of its peers while still staying anchored to the correct ground-truth label. This ensures that the ensemble members learn similar, coherent ways of making predictions.
Gramian Measure
This loss function is used during training to ensure that the internal representations (embeddings) of different PFM experts are spread out across a wide range of possibilities. It encourages the models to cover a low-volume subspace, which helps in detecting genuine data uncertainty rather than just representation mismatch.
Epistemic Uncertainty
This type of uncertainty measures what the ensemble doesn't know because the experts disagree on how to interpret the input. It is formally calculated as the divergence between expert posteriors. High epistemic uncertainty signals a sample where the models are genuinely confused, making it a reliable indicator for decision-making.
Patch-level Consensus
This metric assesses how well different experts agree on where an abnormality is located at the small patch level. It is calculated by averaging the attention coefficients from all experts across that patch. A sharp signal here indicates a highly localized, diagnostically relevant region.

Terminology used across episodes

This episode discusses

The paper

Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning · Read on arXiv

Gbègninougbo Aurel Davy Tchokponhoue, Sevda Ögüt, Ali Idri, Dorina Thanou, Pascal Frossard

UM6P, Ben Guerir, Morocco · EPFL, Lausanne, Switzerland

Pathology foundation models (PFMs) offer generalizable representations for whole-slide image (WSI) analysis, yet their clinical adoption remains limited. Specifically, their predictions lack reliable confidence estimates, and no single PFM is universally best across tasks, which severely undermines trust in medical settings. To overcome this, we propose DICE, a plug-and-play framework that ensembles K frozen PFMs and estimates uncertainty based on their consensus. We align the ensemble members via deep mutual learning and theoretically show that this objective controls an upper bound on epistemic uncertainty. Additionally, we demonstrate that the ensemble localizes abnormalities at the patch level without any explicit supervision. We evaluate DICE on three challenging WSI benchmarks. Notably, our framework provides reliable uncertainty estimates that accurately flag failure-prone cases under in- and out-of-distribution settings, while matching or outperforming SOTA baselines in classification, calibration, and localization. Overall, DICE takes a crucial step toward translating PFMs into uncertainty-aware decision-support systems.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning".

Jane: Pathology foundation models (PFMs) offer generalizable representations for whole-slide image (WSI) analysis,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Welcome back to the show! Today we're looking at some really interesting work coming out of arXiv that’s touching on how we can make sure AI models, especially in complex areas like medical imaging, don't just give answers without telling us how sure they are. We’re talking about the paper "Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning."

Jane: That sounds super practical, Tom; it seems like they are trying to solve a real headache for anyone using these huge foundation models for things like whole-slide image analysis.

Lu: Yeah, the abstract sets up a clear problem: these foundation models give us general representations for whole-slide images, but clinically they struggle because their predictions lack reliable confidence estimates and there isn't one single best model across different tasks <ref:2606.30020#pg1>. That lack of confidence is a huge hurdle when you're dealing with medical decisions.

Meng: From an engineering viewpoint, that uncertainty issue is critical because in a real clinical setting, we absolutely need to know when a model is unsure before we trust its output for diagnosis <ref:2606.30020#pg1>. We can't just let it give us a result without a confidence score attached.

Lalam: I think this paper is really important because it’s about building trust into the AI systems used for diagnosis, which is something we need to focus on as we integrate these tools more deeply into our lives <ref:2606.30020#pg1>. It moves the conversation from just getting an output to understanding the reliability behind that output.

Tom: So they propose DICE, which is this plug-and-play framework that uses an ensemble of K frozen PFMs and models their disagreement as a proxy for uncertainty estimation <ref:2606.30020#pg1>. It seems the central claim here is that the disagreement among these different PFMs acts as a meaningful signal for estimating the actual uncertainty in what they predict.

Jane: That makes sense, Tom; so they align these experts using deep mutual learning to make sure their disagreements actually reflect real uncertainty, and they theoretically show that this alignment process provides an upper bound on the model uncertainty <ref:2606.30020#pg1>.

Lu: That deep mutual learning objective, which involves minimizing a loss term related to KL divergence between expert posteriors, is what establishes that theoretical guarantee about bounding the model uncertainty <ref:2606.30020#pg1>. It’s a solid mathematical underpinning for why this approach should work.

Meng: I wonder about the practical side of running K frozen models and managing all that alignment during training; how feasible is it to do this without making the whole process too slow for real-world use cases <ref:2606.30020#pg1>? We need something deployable, not just a complex research setup.

Lalam: It’s about creating a system where the ensemble itself provides that uncertainty quantification, which I think is a really cool way to make the AI more transparent for users <ref:2606.30020#pg1>. Transparency is key when we're talking about tools that help guide clinical decisions.

Paper summary: Tom: And they don't stop there; they also show that this expert consensus can actually help with localization, which is another quite interesting finding in the paper <ref:2606.30020#pg2>. This goes beyond just knowing *what* something is; it helps them pinpoint *where* it is on the slide.

Jane: I remember reading about how they use an attention map consensus to help pinpoint abnormalities at the patch level without needing any extra supervision for that localization task <ref:2606.30020#pg2>. That ability to localize things automatically is a big deal for diagnostics.

Lu: That localization aspect is significant because it suggests that the agreement among experts isn't just about the final classification label; it translates into a sharp signal for where the pathology actually resides, even without explicit labels for that localization task <ref:2606.30020#pg2>.

Tom: So, to wrap up this summary of "Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning," we see they are shifting away from relying on just one model to provide all the confidence information <ref:2606.30020#pg1>. The authors suggest that by treating disagreements between PFMs as an informative signal, they can translate these models into systems that are aware of their own uncertainty.

Jane: It seems the core message is that by leveraging these diverse AI models and modeling their internal conflicts, we can build a more robust framework for understanding when an AI might be making a guess <ref:2606.30020#pg1>.

Meng: From an engineering perspective, the practical implication here is improving the reliability of whole-slide image analysis by giving us more reliable confidence scores, which directly helps in informed model selection for clinical use <ref:2606.30020#pg1>. That’s something we can measure and deploy.

Lalam: This work really shows how we can leverage diverse AI models to create a more robust and trustworthy diagnostic tool, moving us closer to systems that are honest about their limitations <ref:2606.30020#pg1>.

Tom: Moving on to the conclusion of this discussion regarding "Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning," the title itself really captures the essence of what they achieved <ref:2606.30020#pg1>. The authors framed their work around using deep mutual learning for alignment and focusing squarely on how to estimate uncertainty in pathology models, which is a smart way to frame this complex research.

Jane: It’s about taking these very large foundation models and figuring out exactly how to quantify their uncertainty so that we can actually trust them more in a clinical environment where mistakes have serious consequences <ref:2606.30020#pg1>.

Lu: I think the authors chose that title because it highlights the two main pillars of their work: using deep mutual learning for alignment and focusing on uncertainty estimation, which is a really smart way to frame this research <ref:2606.30020#pg1>. It sets up a very clear roadmap for what they accomplished.

Meng: From my side, I'm thinking about how that title speaks to moving beyond just getting an answer from a model and actually understanding the confidence behind that answer, which is exactly what we need for real-world deployment <ref:2606.30020#pg1>. We need to know when the uncertainty is high enough to flag a case for human review.

Paper summary: Lalam: I feel that title perfectly reflects how they are taking general knowledge from these models and making it trustworthy by adding this layer of uncertainty quantification, which could really help in building better cultural tools based on medical images <ref:2606.30020#pg1>. It’s about making the AI honest.

Tom: Right, and the authors are clearly smart; they've managed to tackle a problem where relying on a single model falls short, showing how ensembling and mutual learning can solve that issue <ref:2606.30020#pg1>. The entire framework is quite clever in how it handles the complexity of multiple experts working together.

Jane: They’ve shown that this approach isn't just a clever trick but is grounded in deep alignment techniques, which makes the resulting uncertainty estimates much more reliable than just picking one model at random <ref:2606.30020#pg1>. The rigor behind the method is what gives us confidence in the results.

Lu: The deep mutual learning objective they use to align the experts really establishes a solid theoretical foundation for how we can contract that uncertainty bound, which is pretty profound research <ref:2606.30020#pg1>. That theoretical proof is what makes this work more than just empirical observation.

Meng: I’m just thinking about the practical implications of that contraction; if the math holds up, it means we have a predictable way to manage model risk when we deploy these systems in a live clinical setting <ref:2606.30020#pg1>. That predictability is what engineers look for most.

Lalam: For me, this implies that as AI becomes more integrated into our daily lives through things like medical diagnostics, having a formal way to say "I'm uncertain about this" could really improve how we interact with and trust those systems in general <ref:2606.30020#pg1>.

Tom: Exactly. So, the main point we’re taking away from "Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning" is that ensembling and mutual learning provide a structured way to quantify uncertainty that goes beyond what any single model can offer <ref:2606.30020#pg1>.

Jane: It’s about taking those powerful foundation models and figuring out how to measure their internal disagreement so we can build systems that are aware of when they might be guessing <ref:2606.30020#pg1>.

Lu: The implication here for the broader field is that we move toward task-agnostic uncertainty quantification, which was something that wasn't really there before because previous methods were often too specific to one particular task or model architecture <ref:2606.30020#pg1>. That’s a big conceptual step forward.

Meng: From an engineering view, it suggests a way to deploy multiple models and use their internal dynamics to gauge reliability dynamically during inference, which is something we can actually start prototyping now <ref:2606.30020#pg1>.

Lalam: If this framework works well, it means we can develop AI tools that are not just accurate but also honest about when they might be guessing, which is a huge step toward real clinical integration and better user interaction <ref:2606.30020#pg1>.

Conclusion: Tom: So we've seen how DICE uses frozen PFMs to estimate uncertainty by modeling their disagreements, and now we're coming to the finish line of this discussion about its title and authors <ref:2606.30020#pg1>.

Jane: Exactly, Tom; the paper is titled "Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning," and it really captures how they take these big foundation models and build a system around them to figure out when they're unsure <ref:2606.30020#pg1>.

Lu: I think the authors chose that title because it highlights the two main pillars of their work: using deep mutual learning for alignment and focusing on uncertainty estimation, which is a really smart way to frame the research <ref:2606.30020#pg1>.

Meng: From my side, I'm thinking about how this title speaks to moving beyond just getting an answer from a model and actually understanding the confidence behind that answer, which is what we need for real-world deployment <ref:2606.30020#pg1>.

Lalam: I feel that title perfectly reflects how they're taking general knowledge and making it trustworthy by adding this layer of uncertainty quantification, which could really help in building better cultural tools based on medical images <ref:2606.30020#pg1>.

Tom: Right, and the authors are clearly smart; they've managed to tackle a problem where single models fall short, showing how ensembling and mutual learning can solve it <ref:2606.30020#pg1>.

Jane: They’ve shown that this approach isn't just a clever trick but is grounded in deep alignment techniques, which makes the uncertainty estimates much more reliable than just guessing <ref:2606.30020#pg1>.

Lu: The deep mutual learning objective they use to align the experts really establishes a solid theoretical foundation for how we can contract that uncertainty bound, which is pretty profound research <ref:2606.30020#pg1>.

Meng: I’m just thinking about the practical implications of that contraction; if the math holds up, it means we have a predictable way to manage model risk when deploying these systems <ref:2606.30020#pg1>.

Lalam: For me, this implies that as AI becomes more integrated into our daily lives through things like medical diagnostics, having a formal way to say "I'm uncertain about this" could really improve how we interact with and trust those systems <ref:2606.30020#pg1>.

Tom: So the core message here is that by using an ensemble and deep learning to model disagreements, DICE gives us a structured way to quantify uncertainty that goes way beyond what any single model can offer <ref:2606.30020#pg1>.

Jane: It’s about taking those powerful foundation models and figuring out exactly how to measure their internal conflicts so we can build systems that are aware of when they might be guessing <ref:2606.30020#pg1>.

Lu: The implication for the broader field is that we move toward task-agnostic uncertainty quantification, which was something that wasn't really there before because previous methods were often too specific to one particular task or model architecture <ref:2606.30020#pg1>.

Meng: From an engineering view, it suggests a way to deploy multiple models and use their internal dynamics to gauge reliability dynamically during inference, which is something we can actually start prototyping now <ref:2606.30020#pg1>.

Lalam: If this framework works well, it means we can develop AI tools that are not just accurate but also honest about when they might be guessing, which is a huge step toward real clinical integration <ref:2606.30020#pg1>.

Tom: This work really shows how ensembling and mutual learning provide a structured way to quantify uncertainty that goes beyond what any single model can offer <ref:2606.30020#pg1>.

Jane: It’s about taking those powerful foundation models and figuring out exactly how to measure their internal conflicts so we can build systems that are aware of when they might be guessing <ref:2606.30020#pg1>.

Lu: The implication here for the broader field is that we move toward task-agnostic uncertainty quantification, which was something that wasn't really there before because previous methods were often too specific to one particular task or model architecture <ref:2606.30020#pg1>.

More episodes

← Home