Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift

arXiv:2607.10358 · cs.CV, cs.AI · Submitted 2026-07-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift".

Jane: The paper was written by Giang Nguyen, Raghav Mehta, Emma A.M. Stanley, Tian Xia, Thi Hao Nguyen et al. from College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam (Affiliation 1) and Imperial College London, United Kingdom (Affiliation 2) and Radiology Department, Vietnam National Cancer Hospital, Hanoi, Vietnam (Affiliation 3) and VinUni-Illinois Smart Health Center, VinUniversity and The Computer Vision and Medical AI Lab, VinUniversity.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: So, to summarize what the authors are doing in "Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift," they've built a massive benchmark. They take fifteen different foundation-model backbones—backbones are essentially the brains of the AI—and test them across three very different medical tasks.

Jane: And those tasks are image-level breast density, exam-level BI-RADS assessment, and binary cancer status classification, which is already complex because clinical definitions vary so much.

Lu: The authors really emphasize that they are testing these models on twelve out-of-distribution or OOD datasets that come from different countries to see if the training data has given them a false sense of security.

Meng: That's the crucial part, because as a practical implementer, I need to know if a model performs well on my local patient population or just in the one dataset it was trained on.

Lalam: This is important for me because it’ shows that relying solely on strong performance in one single source database isn' not enough for robust AI deployment.

Tom: They are using a standardized, or "unified," frozen-backbone linear-probe protocol, which is a very consistent way to evaluate how the representations perform across these diverse settings.

Jane: It seems like they found that while mammography-specific vision-language models—those designed specifically for this type of data—did show strong performance, their general robustness wasn't just due to being trained on breast images.

Lu: That suggests that simply having a lot of pictures of breasts isn't the whole story; you need something more specific about how the model processes clinical information.

Meng: My concern would be whether this specific linear-probe setup is sufficient for real-world deployment, which often requires much deeper fine-tuning than just a simple linear head on frozen features.

Lalam: The goal of this methodology, in my view, is to force us to look at the dataset itself as a primary feature when we evaluate AI performance, not just the model's internal scores.

Improvements: Tom: Looking at the results from "Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift," it seems like there are several key takeaways that suggest improvements in how we approach clinical AI models.

Jane: The authors found that mammography pretraining alone, even if it was on a large dataset, does not consistently improve generalization when facing a domain shift.

Lu: This finding suggests we need to move beyond just focusing on the visual data and start thinking about the broader context of what makes an image useful in different regions.

Meng: From an engineering standpoint, this tells me that if I am building a system for one region, I cannot assume it will work in another region without testing against these types of OOD datasets.

Lalam: It's a call to action for the AI community to prioritize robust validation methods and dataset-level analysis over just achieving high average scores.

Tom: The paper also highlights that even the leading models show heterogeneous performance, meaning they are good at some things but not others depending on which dataset they encounter.

Jane: That heterogeneity is important because it shows that when a model is performing poorly, it's often due to a specific shift in the data structure or acquisition method.

Lu: It’s fascinating to see how feature-space inspection can reveal that even though the AI representation changes based on its training, some useful clinical signals remain intact.

Meng: That helps me understand where my models might be failing, allowing me to pinpoint whether a performance drop is due to the model's internal structure or just the external data quality.

Lalam: We must improve our validation practices so that we are not misled by an average performance score, and instead focus on how the AI behaves under stress.

Conclusions: Tom: As we wrap up "Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift," the authors provide a clear conclusion about what makes a strong clinical AI model.

Jane: They're saying that robust performance isn't just about having seen images; it’s about how the pretraining objective aligns with the specific task and its label origins.

Lu: This is a huge shift in thinking, implying that we need to move towards tailoring our pretraining objectives rather than just applying a standard large-scale model.

Meng: The practical takeaway for me is that if I'm deploying AI in oncology, I have to be very careful about the specific data source and how that data was labeled when I train my foundation model.

Lalam: It's inspiring to see that the future of robust AI requires us to understand not just what the images look like, but how they are structured and categorized across different global contexts.

Tom: The study concludes by emphasizing that OOD evaluation should be the central criterion for assessing clinical utility, not just ID performance.

Jane: It's a sobering reminder that strong source-domain results don' don't guarantee reliable performance when we encounter real-world variability.

Lu: I think the analysis of the UMAP embeddings showing how different views are treated is a powerful illustration of this concept in action.

Meng: I wonder how much more work needs to be done on full fine-tuning, as the linear probe approach in this study is just a baseline representation quality test.

Lalam: We' must ensure that the future AI models will be judged by their reliability when we need them most, not just when they are tested on ideal data.

Wrap-up: Tom: We’ve covered a lot of ground today with "Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift," and I think it's a huge step forward in how we think about clinical AI.

Jane: It's definitely encouraging, Tom, to see the strong performance of tailored models while also being realistic about the challenges when we encounter new datasets.

Lu: The research shows that there are many ways to improve our understanding of how these powerful AI tools by looking at the nuances in dataset-level performance.

Meng: My main point is that this work provides a necessary framework for building practical, deployable, and trustworthy medical AI systems across different regions.

Lalam: We' must use this knowledge to ensure that the advancement of technology serves all patients equally and reliably in diverse settings.

Tom: I think we have covered the core findings well enough for now.

Jane: It’s been a really enlightening discussion on this paper today, and it’s a perfect example of why rigorous benchmarking matters so much, Tom.

Lu: I'm looking forward to seeing how these models are further adapted in subsequent studies.

Meng: And I can't wait to see the practical impact of these findings translated into real-world deployment protocols.

Lalam: Let’s carry this excitement and this rigor with us as we move toward the next paper on our schedule.

Giang Nguyen, Raghav Mehta, Emma A.M. Stanley, Tian Xia, Thi Hao Nguyen, Hieu Pham, Ben Glocker

College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam (Affiliation 1) · Imperial College London, United Kingdom (Affiliation 2) · Radiology Department, Vietnam National Cancer Hospital, Hanoi, Vietnam (Affiliation 3) · VinUni-Illinois Smart Health Center, VinUniversity · The Computer Vision and Medical AI Lab, VinUniversity

cs.CV, cs.AI

Submitted: 2026-07-11

Updated: 2026-08-27

Comments: Accepted at Deep-Brea3th 2026 workshop in conjunction with MICCAI 2026

Code: https://github.com/biomedia-mira/mammo-ood

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: Foundation models (FMs) are increasingly used as image feature extractors for mammography, but their "robustness under external domain shift remains unclear." The utility of an FM depends on whether

Key concepts

Foundation Models
These are large AI models (the 'brains' of the AI) tested in the paper. They are used for various medical tasks, such as image-level breast density assessment or cancer status classification.
Domain Shift/OOD Datasets
This refers to testing a model on data from different sources or countries than the data it was trained on. The authors use these out-of-distribution datasets to test the model's real-world reliability.
Linear-Probe Protocol
A standardized, consistent method used in the study to evaluate how AI representations perform across diverse settings. It involves using a simple linear head on frozen features for evaluation.

Terminology

Summary

Foundation models (FMs) are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remains unclear. The utility of an FM depends on whether its pretrained representation remains useful under distribution shift. This issue is amplified in mammography because of heterogeneous clinical labels: density, BI-RADS, and cancer status differ in granularity, prevalence, and annotation protocols.

To evaluate this generalization, the authors established a benchmark featuring 15 public mammography datasets spanning 12 countries/regions. This benchmark covers three distinct downstream tasks: image-level density, exam-level BI-RADS, and exam-level cancer status. The evaluation utilized a unified frozen-backbone linear-probe protocol, training on source data and testing on 12 task-compatible out-of-distribution (OOD) datasets.

The results show that mammography pretraining alone is insufficient for robust OOD generalization. Instead, OOD performance depends on the pretraining objective, model type (vision-only vs. vision-language), and pretraining dataset diversity.

Key Performance Findings:

  • Mammography-specific vision-language models (Mammo-FM and MaMA) provide the strongest mean OOD performance.

  • DINOv3 remains a competitive vision-only baseline.

  • The leading model is task-dependent rather than universal: Mammo-FM achieves the best BI-RADS OOD AUROC (0.688±0.016), whereas MaMA reaches the best performance for density (0.865±0.006) and cancer status (0.718±0.014).

Analysis of Generalization:

  • VLMs achieve higher mean OOD AUROCs than vision-only models for BI-RADS, density, and cancer status.

  • The authors conclude that mammography language alignment is more important than visual mammography exposure alone.

  • Furthermore, the study found that strong external generalization does not require complete domain invariance, and strong average OOD performance does not necessarily correspond to uniformly strong behavior across OOD datasets.

Feature Space Inspection:

  • Feature-space inspection reveals that useful representations can preserve clinical signal while retaining dataset and acquisition structure.

  • The two strongest mammography-specific VLMs exhibit different qualitative biases: MaMA has the clearest density organization... However, MaMA still preserves visible view separation, whereas Mammo-FM shows better alignment between different views, despite not having any explicit multi-view alignment in its pre-training.

Conclusion:

The benchmark demonstrates that strong source-domain performance and mammography exposure alone do not guarantee robust OOD generalization. The findings suggest that robust mammography representations depend on how pretraining objectives and data sources align with each clinical task and its label provenance. Therefore, the authors assert that mammography foundation models should be assessed using task-specific OOD validation rather than ID performance.

Improvements for AI systems

Improvement: The core evaluation metric for any mammography foundation model must transition from solely assessing In-Distribution (ID) performance to rigorously testing Out-of-Distribution (OOD) generalization across heterogeneous, task-compatible datasets.

What the Improved System Can Do: It ensures that the model's clinical utility is validated not just on training data, but against real-world domain shifts (e.g., different manufacturers, varying acquisition parameters). This prevents deployment of models that perform well in a controlled environment but fail clinically due to lack of generalization.

Improvement: Prioritize the adoption and fine-tuning of dedicated mammography VLMs (e.g., architectures akin to MaMA or Mammo-FM), rather than relying on generalized natural image Self-Supervised Learning (SSL).

What the Improved System Can Do: The model achieves superior OOD performance in clinical tasks (BI-RADS and Cancer Status) because it integrates semantic understanding derived from medical reports/metadata, not just visual features. It can robustly handle unseen datasets by correlating visual cues with known clinical outcomes.

Improvement: Implement pretraining objectives that leverage specific clinical information beyond standard image-text matching. This involves two distinct strategies based on the target task:

  • For Density/Visual Cues (MaMA approach): Integrate tabular metadata (e.g., density, BI-RADS scores) into the language conditioning during pretraining to ensure explicit correlation between tissue composition and visual patterns.

  • For Clinical Outcomes (Mammo-FM approach): Utilize actual clinical diagnostic reports as the primary supervisory signal during pretraining, ensuring the model aligns its feature space with established medical terminology and final diagnoses.

What the Improved System Can Do: The system becomes highly specialized. It can predict density/texture patterns with high fidelity (MaMA) or provide a clinically actionable diagnosis that matches expert reporting standards (Mammo-FM), maximizing its utility for specific clinical workflows.

Improvement: Incorporate periodic UMAP visualization and feature-space analysis into the continuous integration/validation pipeline for all five representative backbones (DINOv3, MaMA, etc.).

What the Improved System Can Do: The system can proactively identify model degradation or drift by monitoring how the embedding geometry changes relative to manufacturer and view position. It ensures that strong generalization is accompanied by a stable, meaningful internal representation (i.e., verifying that MaMA’s density-related clusters remain cohesive).

Improvement: Implement strict regularization and constraint measures to prevent the single-source adaptation of a model on one dataset (e.g., RSNA-site2) must be insufficient for deployment. The system must treat localized, single-dataset training as a failure mode.

What the Improved System Can Do: It prevents deployment of models that appear highly successful only on the source data, ensuring that the operational model is robust across all 12 countries/regions identified in the benchmark.

Abstract

Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remains unclear. We benchmark 15 foundation-model backbones across breast density, BI-RADS severity, and cancer status using a unified frozen-backbone linear-probe protocol, training on 3 source datasets and evaluating on 12 task-compatible out-of-distribution (OOD) datasets after label harmonization. Mammography-specific vision-language models (Mammo-FM and MaMA) provide the strongest mean OOD performance, but robustness is not explained by mammography exposure alone. DINOv3 remains a competitive vision-only baseline, and mammography-adapted pretraining does not consistently improve generalization. Dataset-level analysis further shows that even leading models show heterogeneous performance across datasets. Feature-space inspection reveals that useful representations can preserve clinical signal while retaining dataset and acquisition structure. These findings highlight dataset-level OOD evaluation as a central criterion for assessing mammography representations. Our code is publicly available: https://github.com/biomedia-mira/mammo-ood.

Sources

Related papers