Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift

summary

Video file (mp4)

The gist

Foundation models (FMs) are increasingly used as image feature extractors for mammography, but their "robustness under external domain shift remains unclear." The utility of an FM depends on whether

In short

The episode discusses 'Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift.' Hosts analyze how foundation models perform across diverse, out-of-distribution datasets from different countries. The conclusion emphasizes that robust clinical AI requires evaluation beyond single datasets, focusing on dataset structure and task alignment.

Key concepts

Foundation Models
These are large AI models (the 'brains' of the AI) tested in the paper. They are used for various medical tasks, such as image-level breast density assessment or cancer status classification.
Domain Shift/OOD Datasets
This refers to testing a model on data from different sources or countries than the data it was trained on. The authors use these out-of-distribution datasets to test the model's real-world reliability.
Linear-Probe Protocol
A standardized, consistent method used in the study to evaluate how AI representations perform across diverse settings. It involves using a simple linear head on frozen features for evaluation.

Terminology used across episodes

This episode discusses

The paper

Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift · Read on arXiv

Giang Nguyen, Raghav Mehta, Emma A.M. Stanley, Tian Xia, Thi Hao Nguyen, Hieu Pham, Ben Glocker

College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam (Affiliation 1) · Imperial College London, United Kingdom (Affiliation 2) · Radiology Department, Vietnam National Cancer Hospital, Hanoi, Vietnam (Affiliation 3) · VinUni-Illinois Smart Health Center, VinUniversity · The Computer Vision and Medical AI Lab, VinUniversity

Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remains unclear. We benchmark 15 foundation-model backbones across breast density, BI-RADS severity, and cancer status using a unified frozen-backbone linear-probe protocol, training on 3 source datasets and evaluating on 12 task-compatible out-of-distribution (OOD) datasets after label harmonization. Mammography-specific vision-language models (Mammo-FM and MaMA) provide the strongest mean OOD performance, but robustness is not explained by mammography exposure alone. DINOv3 remains a competitive vision-only baseline, and mammography-adapted pretraining does not consistently improve generalization. Dataset-level analysis further shows that even leading models show heterogeneous performance across datasets. Feature-space inspection reveals that useful representations can preserve clinical signal while retaining dataset and acquisition structure. These findings highlight dataset-level OOD evaluation as a central criterion for assessing mammography representations. Our code is publicly available: https://github.com/biomedia-mira/mammo-ood.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift".

Jane: The paper was written by Giang Nguyen, Raghav Mehta, Emma A.M. Stanley, Tian Xia, Thi Hao Nguyen et al. from College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam (Affiliation 1) and Imperial College London, United Kingdom (Affiliation 2) and Radiology Department, Vietnam National Cancer Hospital, Hanoi, Vietnam (Affiliation 3) and VinUni-Illinois Smart Health Center, VinUniversity and The Computer Vision and Medical AI Lab, VinUniversity.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: So, to summarize what the authors are doing in "Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift," they've built a massive benchmark. They take fifteen different foundation-model backbones—backbones are essentially the brains of the AI—and test them across three very different medical tasks.

Jane: And those tasks are image-level breast density, exam-level BI-RADS assessment, and binary cancer status classification, which is already complex because clinical definitions vary so much.

Lu: The authors really emphasize that they are testing these models on twelve out-of-distribution or OOD datasets that come from different countries to see if the training data has given them a false sense of security.

Meng: That's the crucial part, because as a practical implementer, I need to know if a model performs well on my local patient population or just in the one dataset it was trained on.

Lalam: This is important for me because it’ shows that relying solely on strong performance in one single source database isn' not enough for robust AI deployment.

Tom: They are using a standardized, or "unified," frozen-backbone linear-probe protocol, which is a very consistent way to evaluate how the representations perform across these diverse settings.

Jane: It seems like they found that while mammography-specific vision-language models—those designed specifically for this type of data—did show strong performance, their general robustness wasn't just due to being trained on breast images.

Lu: That suggests that simply having a lot of pictures of breasts isn't the whole story; you need something more specific about how the model processes clinical information.

Meng: My concern would be whether this specific linear-probe setup is sufficient for real-world deployment, which often requires much deeper fine-tuning than just a simple linear head on frozen features.

Lalam: The goal of this methodology, in my view, is to force us to look at the dataset itself as a primary feature when we evaluate AI performance, not just the model's internal scores.

Improvements: Tom: Looking at the results from "Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift," it seems like there are several key takeaways that suggest improvements in how we approach clinical AI models.

Jane: The authors found that mammography pretraining alone, even if it was on a large dataset, does not consistently improve generalization when facing a domain shift.

Lu: This finding suggests we need to move beyond just focusing on the visual data and start thinking about the broader context of what makes an image useful in different regions.

Meng: From an engineering standpoint, this tells me that if I am building a system for one region, I cannot assume it will work in another region without testing against these types of OOD datasets.

Lalam: It's a call to action for the AI community to prioritize robust validation methods and dataset-level analysis over just achieving high average scores.

Tom: The paper also highlights that even the leading models show heterogeneous performance, meaning they are good at some things but not others depending on which dataset they encounter.

Jane: That heterogeneity is important because it shows that when a model is performing poorly, it's often due to a specific shift in the data structure or acquisition method.

Lu: It’s fascinating to see how feature-space inspection can reveal that even though the AI representation changes based on its training, some useful clinical signals remain intact.

Meng: That helps me understand where my models might be failing, allowing me to pinpoint whether a performance drop is due to the model's internal structure or just the external data quality.

Lalam: We must improve our validation practices so that we are not misled by an average performance score, and instead focus on how the AI behaves under stress.

Conclusions: Tom: As we wrap up "Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift," the authors provide a clear conclusion about what makes a strong clinical AI model.

Jane: They're saying that robust performance isn't just about having seen images; it’s about how the pretraining objective aligns with the specific task and its label origins.

Lu: This is a huge shift in thinking, implying that we need to move towards tailoring our pretraining objectives rather than just applying a standard large-scale model.

Meng: The practical takeaway for me is that if I'm deploying AI in oncology, I have to be very careful about the specific data source and how that data was labeled when I train my foundation model.

Lalam: It's inspiring to see that the future of robust AI requires us to understand not just what the images look like, but how they are structured and categorized across different global contexts.

Tom: The study concludes by emphasizing that OOD evaluation should be the central criterion for assessing clinical utility, not just ID performance.

Jane: It's a sobering reminder that strong source-domain results don' don't guarantee reliable performance when we encounter real-world variability.

Lu: I think the analysis of the UMAP embeddings showing how different views are treated is a powerful illustration of this concept in action.

Meng: I wonder how much more work needs to be done on full fine-tuning, as the linear probe approach in this study is just a baseline representation quality test.

Lalam: We' must ensure that the future AI models will be judged by their reliability when we need them most, not just when they are tested on ideal data.

Wrap-up: Tom: We’ve covered a lot of ground today with "Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift," and I think it's a huge step forward in how we think about clinical AI.

Jane: It's definitely encouraging, Tom, to see the strong performance of tailored models while also being realistic about the challenges when we encounter new datasets.

Lu: The research shows that there are many ways to improve our understanding of how these powerful AI tools by looking at the nuances in dataset-level performance.

Meng: My main point is that this work provides a necessary framework for building practical, deployable, and trustworthy medical AI systems across different regions.

Lalam: We' must use this knowledge to ensure that the advancement of technology serves all patients equally and reliably in diverse settings.

Tom: I think we have covered the core findings well enough for now.

Jane: It’s been a really enlightening discussion on this paper today, and it’s a perfect example of why rigorous benchmarking matters so much, Tom.

Lu: I'm looking forward to seeing how these models are further adapted in subsequent studies.

Meng: And I can't wait to see the practical impact of these findings translated into real-world deployment protocols.

Lalam: Let’s carry this excitement and this rigor with us as we move toward the next paper on our schedule.

More episodes

← Home