Optimizing Breast Cancer Detection in Mammograms: A Comprehensive Study of Transfer Learning, Resolution Reduction, and Multi-View Classification

arXiv:2503.19945 · eess.IV, cs.AI, cs.CV · Submitted 2025-03-25 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Optimizing Breast Cancer Detection in Mammograms".

Jane: Mammography remains central to early breast cancer detection, yet interpreting mammograms requires expertise and traditional methods have limitations in accuracy <ref:2503.19945/pg1>.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: This paper focuses on "Optimizing Breast Cancer Detection in Mammograms: A Comprehensive Study of Transfer Learning, Resolution Reduction, and Multi-View Classification." We’re seeing a lot of investigation into how different AI techniques interact with the specific challenges of mammography.

Jane: The authors are Daniel G. P. Petrini and Hae Yong Kim from the Department of Electronic Systems Engineering at Polytechnic School in University of São Paulo. They are trying to find better ways to use deep learning for this imaging technique because interpreting mammograms still requires a lot of expertise, even with current computer-aided detection systems one.

Lu: The core message is that they are systematically testing five specific research questions, which cover everything from the role of patch classifiers to how well we can transfer knowledge from natural images to mammography.

Meng: They also look at things like using "learn-to-resize" instead of just standard shrinking, and how combining multiple views actually helps the classification accuracy. That’s a lot of moving parts in one study.

Lalam: I think the core idea is that simply applying a standard deep learning model isn't enough for mammography because the images are so specialized, and this paper tries to systematically find out the best path forward by looking at all these different strategies #pg0.

The paper's summary: Tom: The summary of this study shows they compared different classification methods, checking patch classifiers against direct classifiers, and they also tested how much performance improves when you use multiple views instead of just one.

Jane: It boils down to this: single-view methods have limitations, but integrating two or more views seems to give a statistically better result than just averaging the results from separate single-view classifications.

Lu: They found that for instance, on the VinDrMammo dataset, the best two-view classifier beat both simple averaging and maximizing methods when you look at statistical tests like p=zero point zero zero three zero and p=zero point zero two eight six #pg31. That’s a solid finding about feature fusion.

Meng: So, if we were building a system right now, this suggests that trying to get two different images—say, one from the top and one from the side—and feeding them into the AI at once is definitely a better approach than running two separate classifiers and stitching their scores together.

Lalam: It means that by combining features from both views, we get extra information about where lesions are located spatially, which simple averaging just can’t capture on its own #pg17.

The paper's improvements: Tom: Now the paper points out some specific areas where they think things could be better. They suggest that patch-based pretraining isn't actually the best starting point for lower-quality mammography datasets like CBIS-DDSM.

Jane: That’s a practical improvement because it simplifies your pipeline; if patch training doesn't help much, you should just go straight to using weights already trained on standard natural images, which is what they call direct classification.

Lu: They also look at image resolution reduction and found that the "learn-to-resize" technique didn't perform as well as conventional fixed resizing when dealing with mammograms specifically.

Meng: That’s interesting because we often think machine learning can fix any downsampling issue, but this study shows that for these specific X-ray images, just sticking to a standard fixed resizing method might be more reliable than using complex learned techniques.

Lalam: And they actually conclude that reducing the resolution of high-quality digital mammograms before classifying them isn't recommended because it caused a significant drop in performance compared to the original one thousand one hundred fifty-two times eight hundred ninety-two resolution #pg18.

Conclusion: Tom: So, wrapping up on this "Optimizing Breast Cancer Detection in Mammograms: A Comprehensive Study of Transfer Learning, Resolution Reduction, and Multi-View Classification," the main message is that multi-view classification substantially outperformed single-view classifiers on both datasets.

Jane: And there’s a positive correlation found between how well a model does natural image tasks and its performance on mammograms when patch pretraining is involved, with numbers like r = zero point six four nine one for CBIS-DDSM #pg19.

Lu: This suggests that for high-quality data, using a modern base model combined with the patch-based pretraining approach is the way to go, which points toward prioritizing backbones like ConvNeXt or DenseNet169 #pg2.

Meng: For practical deployment, this means we should focus our development efforts on building systems that leverage those two-view strategies because they give us that extra spatial information we talked about earlier, which is much more useful for diagnosis #pg6.

Lalam: I think the overall implication is a strong push toward multi-view strategies in future mammogram analysis frameworks because the performance gains are clear and statistically significant #pg6.

Tom: That's a lot to chew on with this paper on "Optimizing Breast Cancer Detection in Mammograms: A Comprehensive Study of Transfer Learning, Resolution Reduction, and Multi-View Classification." We’ve seen how moving to multi-view classification really gives us that spatial context.

Jane: It definitely shows that combining different views is statistically better than just mixing the results from separate single-view classifiers.

Lu: And we should keep an eye on those findings about using patch pretraining for high-quality data, which points toward ConvNeXt or DenseNet169 as good backbone choices.

Meng: From an engineering perspective, it’s clear that prioritizing those two-view strategies is the way to go because they deliver that extra spatial information needed for better diagnosis #pg6.

Lalam: Ultimately, this paper gives us a strong push toward multi-view strategies in future mammogram analysis frameworks because the performance gains are clear and statistically significant #pg6.

Daniel G. P. Petrini, Hae Yong Kim

Department of Electronic Systems Engineering, Polytechnic School, University of São Paulo

eess.IV, cs.AI, cs.CV

Submitted: 2025-03-25

Updated: 2026-10-03

Code: https://github.com/dpetrini/multiple-view

Importance score: 83/100

The gist: Mammography remains central to early breast cancer detection, yet interpreting mammograms requires expertise and traditional methods have limitations in accuracy <ref:2503.19945#pg2>.

Key concepts

Patch Classifier (PBC)
This method involves training a classifier on small patches of an image rather than the whole image. The resulting model's weights are then used for classification. The study found this approach was not advantageous for lower-quality mammography data compared to using ImageNet pre-trained weights.
Learn-to-Resize (LRC)
This is a machine learning technique used to resize images during processing, as opposed to fixed resizing. While effective for natural images, the study concluded that LRC underperformed conventional fixed resizing on mammograms and was not recommended for high-quality digital mammograms.
Multi-View Classification
This involves using two or more separate views of a mammogram to make a diagnosis. The research demonstrated that classifying two views simultaneously is statistically superior to processing individual views separately and combining their results using simple methods like averaging or taking the maximum value.

Terminology

Summary

Mammography remains central to early breast cancer detection, yet interpreting mammograms requires expertise and traditional methods have limitations in accuracy <ref:2503.19945#pg2>. This study systematically investigates key research questions concerning patch classifiers, transferability of natural-image backbones, the advantages of learn-to-resize, multi-view integration, and robustness across varying image qualities to establish new state-of-the-art benchmarks for breast cancer screening tools <ref:2503.19945#pg2>.

How it works

The research addresses five key questions influencing CNN classification: (1) training strategies (e.g., patch-based pretraining versus end-to-end learning), (2) choice of backbone architecture, (3) image resolution and downsampling techniques, (4) single-view versus multi-view integration, and (5) model performance on datasets with varying image quality <ref:2503.19945#pg3>. The methodology involves exploring an evolutionary scale of base models starting from ResNet [27] up to ConvNeXt [37], and evaluating different training approaches such as patch-based-pretrain, resizing, and others <ref:2503.19945#pg9>.

Single-View Classification Strategies

The study compares two primary initialization methods for single-view classifiers: the Patch Classifier (PBC) and the Direct Classifier (DC). The PBC involves training a patch classifier and leveraging its weights, while the DC applies transfer learning directly from ImageNet pre-trained weights <ref:2503.19945#pg10>. For the CBIS-DDSM dataset, both approaches yielded comparable performance when implemented with the EfficientNet-B3 backbone, with PBC achieving an AUC of 0.8325±0.0171 and DC achieving 0.8313±0.0172 <ref:2503.19945#pg13>. The analysis suggests that pretraining on patches is not a suitable strategy for building classifiers for lower-quality mammography datasets such as CBIS-DDSM because the patch-based classifier does not offer a significant advantage over directly utilizing ImageNet-pretrained weights <ref:2503.19945#pg13>.

Resolution Reduction Techniques

The investigation into image resolution explored reducing input size by half, resulting in 576×448 pixels, using either conventional interpolation (Fixed Resizing Classifier or FRC) or a machine learning-based technique called learn-to-resize (LRC) <ref:2503.19945#pg14>. The results indicated that the Learn to Resize technique underperformed compared to fixed resizing, suggesting that while LRC is effective for natural images, it is not well-suited for mammograms <ref:2503.19945#pg18>. Furthermore, downscaling the input images led to a significant decrease in AUC when compared to the original 1152×892 resolution, leading to the conclusion that it is not recommended to reduce the resolution of high-quality digital mammograms before classifying them <ref:2503.19945#pg19>.

Multi-View Classification and Robustness

The contribution of multi-view integration was demonstrated by comparing single-view versus two-view classifiers on both datasets. For the CBIS-DDSM dataset, the best two-view classifier achieved an AUC of 0.8643, which surpasses the AUC of 0.8418 reported in [6] using EfficientNet-B0 <ref:2503.19945#pg15>. When comparing the two-view model against processing individual views and combining their results using mean or maximum operations, statistical tests (DeLong test) showed that classifying two views simultaneously is statistically superior to classifying individual views and combining their results using mean or maximum <ref:2503.19945#pg15>. This superiority was confirmed on the VinDr-Mammo dataset, where the two-view PBC classifier outperformed both averaging and maximizing approaches, with p-values of p = 0.0030 and p = 0.0025.

Conclusions on Optimal Strategy

The study concludes that employing a two-view classification approach substantially outperformed single-view classifiers across both datasets <ref:2503.19945#pg20>. A positive correlation emerged between a model’s performance on natural-image classification tasks and mammogram classification when patch-based pretraining was applied (CBIS-DDSM: r = 0.6491; VinDr-Mammo: r = 0.7288), whereas without patch-based pretraining, this correlation was negligible or negative <ref:2503.19945#pg20>. Ultimately, the findings strongly support the adoption of multi-view strategies in future mammogram analysis frameworks <ref:2503.19945#pg20>.

REFERENCES

[1] World Cancer Research Fund International WCRF. World cancer research fund international, 2024. Available at: https://www.wcrf.org/dietandcancer/worldwidecancer-data/. Accessed on: Sept. 12, 2024.<ref:2503.19945#pg2>

[3] Thijs Kooi, Geert Litjens, Bram Van Ginneken, Albert Gubern-Mérida, Clara I Sánchez, Ritse Mann, Ard den Heeten, and Nico Karssemeijer. Large scale deep learning for computer aided detection of mammographic lesions. Medical image analysis, 35:303–312, 2017.<ref:2503.19945#pg8>

[4] L. Shen, L.R. Margolies, J. H. Rothstein, E. Fluder, R. McBride, and W. Sieh. Deep learning to improve breast cancer detection on screening mammography. Scientific Reports, 9(1):1–12, 2019.<ref:2503.19945#pg5>

[5] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009.<ref:2503.19945#pg2>

[6] D. G. P. Petrini, C. Shimizu, R. A. Roela, G. V. Valente, M. A. A. K Folgueira, and H Y Kim Breast cancer diagnosis in two-view mammography using end-to-end trained efficientnet-based convolutional network IEEE Access, 10:77723–77731, 2022.<ref:2503.19945#pg7>

[8] S. M. McKinney, M. Sieniek, V. Godbole, et al. International evaluation of an ai system for breast cancer screening Nature, 577(7788):89–94, 2020.<ref:2503.19945#pg2>

[9] K. et al. Bowyer. The digital database for screening mammography. In: Third international workshop on digital mammography, page 27, 1996.<ref:2503.19945#pg4>

[10] P SUCKLING J. The mammographic image analysis society digital mammogram database. Digital Mammo, pages 375–386, 1994.<ref:2503.19945#pg4>

[11] Thomas M Lehmann, Mark O Gúld, Christian Thies, Bartosz Plodowski, Daniel Keysers, Bastian Ott, and Henning Schubert. Irma–content-based image retrieval in medical applications. In MEDINFO 2004, pages 842–846. IOS Press, 2004.<ref:2503.19945#pg4>

[12] Inês C Moreira, Igor Amaral, Inês Domingues, António Cardoso, Maria Joao Cardoso, and Jaime S Cardoso. Inbreast: toward a full-field digital mammographic database. Academic radiology, 19(2):236–248, 2012.<ref:2503.19945#pg5>

[13] R. S. Lee, F. Gimenez, A Hoogi, K Miyake, M Gorovoy, and D L Rubin. A curated mammography data set for use in computer-aided detection and diagnosis research. Scientific Data, 4(1):1–9, 2017.<ref:2503.

Improvements for AI systems

  1. textbfTransfer Learning Strategy Refinement for Low-Quality Data (CBIS-DDSM): The paper concludes that patch-based pretraining is not a suitable strategy for building classifiers for lower-quality mammography datasets such as CBIS-DDSM. This suggests shifting from the Patch Classifier (PBC) approach to the Direct Classifier (DC) approach, which showed better performance on this dataset, thereby simplifying the training pipeline while maintaining high accuracy.

  2. textbfModel Backbone Selection Based on Pretraining Correlation: The analysis shows a positive correlation emerged between a model’s performance on natural-image classification tasks and mammogram classification when patch-based pretraining was applied (CBIS-DDSM: r = 0.6491; VinDr-Mammo: r = 0.7288). This indicates that for high-quality data (VinDr-Mammo), combining a modern base model combined with the PBC approach is the optimal strategy, suggesting prioritizing state-of-the-art backbones like ConvNeXt or DenseNet169.

  3. textbf Multi-View Integration for Superior Diagnosis: The study demonstrates that classifying two views simultaneously is statistically superior to classifying individual views and combining their results using mean or maximum operations on both datasets (VinDr-Mammo: p = 0.0030, Mean; p = 0.0286, Maximum). This enables a system that fuses CC and MLO features jointly to provide a substantially outperformed single-view classifiers across both datasets.

  4. textbf Resolution Management Strategy for Digital Mammograms: The findings indicate that reducing the resolution of high-quality digital mammograms before classifying them is not recommended, as the best classifier without resizing performed better than those with fixed or learned resizing, suggesting that using an even higher resolution (e.g., 2304×1792) is likely to improve performance.

  5. textbf Enhanced Two-View Feature Fusion: The two-view classifier achieves high performance by combining features from both views, which provides additional information, such as the spatial locations of lesions, which is not captured by simply aggregating the outputs of separate views. This allows for a more nuanced interpretation of lesion location and characteristics compared to simple averaging or maximization.

Sources

Related papers