Transfer Learning and Machine Learning for Training Five Year Survival Prognostic Models in Early Breast Cancer

arXiv:2509.23268 · cs.LG, cs.CY · Submitted 2025-09-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Transfer Learning and Machine Learning for Training Five Year Survival Prognostic Models in Early Breast Cancer".

Jane: The paper was written by Lisa Pilgram, Kai Yang, Ana-Alicia Beltran-Bless, Gregory R. Pond, Lisa Vandermeer et al. from University of Ottawa and Children's Hospital of Eastern Ontario Research Institute and Charité - Universitaetsmedizin Berlin and McMaster University and The Ottawa Hospital Research Institute and Université de Montréal and Queen's University and University of Edinburgh and Ontario Institute for Cancer Research and University of Toronto and Leiden University Medical Center and St. Augustinus Hospital and University Center Mainz and National and Kapodistrian University of Athens.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. We've got a fascinating paper on the table today, and it's all about breast cancer survival prediction. The title is "Transfer Learning and Machine Learning for Training Five Year Survival Prognostic Models in Early Breast Cancer." Jane, I have to say, when I first saw this title, I thought, okay, this is a technical methods paper, but the more I dig in, the more I realize how much this could actually change clinical practice.

Jane: Absolutely, Tom. And for our listeners who might not be deep in the weeds of machine learning, let's break down what that title actually means. "Transfer learning" is essentially taking a model that's already been trained on a huge amount of data, and then fine-tuning it on a smaller, more specific dataset. Think of it like learning to drive a car, and then learning to drive a truck — you don't start from scratch, you build on what you already know.

Tom: Right, and that's exactly what they did here. They took an existing, widely-used prognostic tool called PREDICT v3, which estimates survival for breast cancer patients, and they fine-tuned it on data from a clinical trial called MA.twenty-seven. The idea being, can we make this general tool better by adapting it to a specific patient population?

Jane: And that's the key question, because these pre-trained models are built on one population, but the patients you see in your clinic might be different. The authors are essentially asking, can we bridge that gap? And the answer, based on what we're seeing, is a resounding yes, at least in some cases.

Tom: Yeah, the results are pretty striking. The fine-tuned model, they call it f-PREDICT v3, saw its calibration error drop from zero point zero four two down to zero point zero zero five. For the non-experts, that means the model's predicted survival probabilities got much, much closer to what actually happened to the patients.

Jane: And that's not just a statistical nicety. If you're a doctor telling a patient, "you have an eighty percent chance of surviving five years," you want that number to be accurate. A miscalibrated model could lead to over-treatment or under-treatment, which is a huge deal.

Tom: Exactly. And what's interesting is that they didn't just stop at fine-tuning. They also trained completely new models from scratch on the MA.twenty-seven data, using something called Random Survival Forests and XGBoost, which are pretty standard machine learning tools. And they compared all of these against the original, un-tuned PREDICT v3.

Jane: So we've got a three-way comparison happening here: the old model, the fine-tuned model, and the brand-new models. And the results, as we'll get into in a bit, are not as simple as "one size fits all." There are trade-offs, especially when you start looking at external validation on other datasets.

Tom: Right, and that's where it gets really interesting, because a model that works great on one population might fall apart on another. We'll get into that in the next segment, but for now, let's just say this paper is asking the right questions about how we build and adapt these tools.

Jane: And it's a question that matters for so many diseases, not just breast cancer. If we can figure out how to adapt these models effectively, we can make personalized medicine more accessible and more accurate for everyone.

Tom: Stay tuned, because we're about to dig into the actual methods and results. You don't want to miss this.

Summary: Tom: Welcome back. We're diving deeper into "Transfer Learning and Machine Learning for Training Five Year Survival Prognostic Models in Early Breast Cancer." Last segment, we talked about the core idea of fine-tuning a pre-trained model. Now, let's get into what the paper actually did and found.

Jane: Right, and the first thing that stands out is the data. They used the MA.twenty-seven trial, which had over seven thousand five hundred postmenopausal women with early-stage, hormone receptor-positive breast cancer. That's a very specific group, and it's a great test case because it's a real clinical trial with real outcomes.

Tom: And the outcome they were predicting is five-year survival, which is a key milestone in breast cancer. But here's the thing — the dataset was heavily imbalanced. Only one hundred eighty-seven patients, about two point five percent, actually had a breast cancer-related death within that five-year window. So the model has to learn from a very small number of events.

Jane: That's a classic problem in survival analysis, and it's one of the reasons why they considered re-balancing the data. They tested a technique called ROSE, which essentially creates synthetic examples of the minority class to balance things out. But interestingly, it didn't help. In fact, it made calibration much worse.

Tom: Yeah, that was a surprising negative result. The calibration error went from something like zero point zero zero five up to zero point three or higher in some models. So they wisely dropped that approach and just trained on the original, imbalanced data.

Jane: And that's a good lesson for anyone doing this kind of work — sometimes the standard tricks don't work, and you have to trust the data as it is. But the bigger story here is the comparison between the models. The fine-tuned PREDICT v3 and the Random Survival Forest both performed really well in terms of calibration, with ICI values around zero point zero zero three to zero point zero zero five.

Tom: For our listeners, that ICI is the Integrated Calibration Index, and lower is better. A value of zero point zero zero five means the average difference between predicted and observed survival is half a percent. That's incredibly accurate.

Jane: And discrimination, which is the model's ability to tell who will survive and who won't, was also solid. The AUC values ranged from about zero point seven four to zero point eight zero. The fine-tuned model actually had the best discrimination at zero point seven nine nine, which is a meaningful improvement over the original PREDICT v3 at zero point seven three eight.

Tom: But here's where it gets complicated. They also did external validation on two other datasets — SEER, which is a big US cancer registry, and TEAM, which is another clinical trial. And the results were mixed. On SEER, the fine-tuned model and the others still looked good. On TEAM, they didn't.

Jane: That's the reality check. The MA.twenty-seven-optimized models actually performed worse on TEAM than the original PREDICT v3. The AUC dropped, and the calibration got worse. So the improvements they saw in MA.twenty-seven didn't fully generalize.

Tom: And that's a crucial finding, because it tells us that transfer learning isn't a magic bullet. It works when the new population is similar to the training population, but when there's a big shift, it can actually hurt. We'll talk more about what that means for practice in the next segment.

Jane: For now, the takeaway is that this paper is a rigorous, honest look at what works and what doesn't when you're trying to build these prognostic models. It's not just a success story — it's a nuanced picture of the challenges.

Improvements: Tom: Welcome back to our discussion of "Transfer Learning and Machine Learning for Training Five Year Survival Prognostic Models in Early Breast Cancer." We've covered the basics and the mixed results on external validation. Now, let's talk about what this paper actually suggests we should do differently.

Jane: Right, and one of the most practical improvements is handling missing data. The original PREDICT v3 simply couldn't generate a prediction for about a quarter of the MA.twenty-seven patients — between twenty-three point eight percent and twenty-five point eight percent — because key information like tumor grade or size was missing.

Tom: That's a huge gap. If you're a clinician and the tool just says "sorry, I can't help you" for one in four patients, that's a problem. But the machine learning models, like the Random Survival Forest and the ensemble, could handle missing data internally. They could predict survival for every single patient.

Jane: And that's a real improvement. The tree-based models use something called surrogate splits, which essentially find alternative ways to split the data when a variable is missing. So they don't just throw up their hands — they work with what they have.

Tom: And the ensemble, which combines the fine-tuned PREDICT, the Random Survival Forest, and XGBoost, was designed to fall back on the machine learning models when PREDICT couldn't make a prediction. So you get the best of both worlds — the accuracy of the fine-tuned model when it works, and the robustness of the ML models when it doesn't.

Jane: That's a really smart design. And the paper also looked at which variables matter most for survival prediction. Using SHAP analysis, they found that patient age, nodal status, tumor grade, and tumor size were consistently the most important factors across all models.

Tom: That makes sense clinically. Those are the classic prognostic factors that oncologists have been using for decades. But it's nice to see that the machine learning models are picking up on the same signals, which gives us confidence that they're learning something real.

Jane: And interestingly, treatment variables like chemotherapy and radiotherapy were ranked lower in importance. That doesn't mean they don't matter, but it suggests that for predicting overall survival, the tumor biology and patient characteristics dominate.

Tom: Now, the paper also suggests that if you're going to use these models in a new population, you need to be careful. The external validation on TEAM showed that the MA.twenty-seven-tuned models didn't transfer well. So the improvement here is really about knowing when to trust your model and when to re-tune it.

Jane: And that's a practical takeaway for anyone building these tools. You can't just train once and deploy everywhere. You need to validate on your own population, and if it doesn't work, you need to fine-tune again.

Tom: Exactly. And the authors are pretty clear that this isn't a one-size-fits-all solution. But the framework they've laid out — fine-tuning, ensemble integration, and careful external validation — is a solid blueprint for building better prognostic models.

Jane: And that's the real contribution here. It's not just about breast cancer. It's about how we should approach model development and adaptation in medicine more broadly.

Conclusion: Tom: And that brings us to the end of our discussion on "Transfer Learning and Machine Learning for Training Five Year Survival Prognostic Models in Early Breast Cancer." Jane, what's the big picture here?

Jane: The big picture, Tom, is that this paper gives us a realistic and practical roadmap for improving survival prediction in breast cancer. Fine-tuning an existing tool like PREDICT v3 can dramatically improve its accuracy on a new population, and machine learning models can fill in the gaps when data is missing.

Tom: And the ensemble approach — combining the fine-tuned model with the ML models — gives you a robust tool that can handle real-world messiness. But the external validation results remind us that these models aren't universal. They need to be tested and adapted for each new setting.

Jane: Right. The paper is honest about the limitations. The TEAM dataset didn't show the same benefits, which tells us that dataset shift is a real challenge. But the framework they've built — train, fine-tune, ensemble, validate — is exactly what we need to move forward.

Tom: And for patients, this could mean more accurate survival estimates, which can guide treatment decisions and reduce both over-treatment and under-treatment. That's a meaningful impact.

Jane: Absolutely. And I think the biggest takeaway is that we have the tools to make these models better, but we have to be thoughtful about how we use them. It's not just about throwing more data at a model — it's about adapting and validating.

Tom: Well said, Jane. We've covered a lot of ground today, from the basics of transfer learning to the nuances of external validation. I hope our listeners found this as fascinating as we did.

Jane: And we're already looking forward to the next paper. Thanks for tuning in, everyone. We'll see you next time.

Tom: Take care, and keep questioning the numbers.

Lisa Pilgram, Kai Yang, Ana-Alicia Beltran-Bless, Gregory R. Pond, Lisa Vandermeer, John Hilton, Marie-France Savard, Andréanne Leblanc, Lois Sheperd, Bingshu E. Chen, John M. S. Bartlett, Karen J. Taylor, Jane Bayani, Sarah L. Barker, Melanie Spears, Cornelis J. H. van der Velde, Elma Meershoek-Klein Kranenbarg, Luc Dirix, Elizabeth Mallon, Annette Hasenburg, Christos Markopoulos, Lamin Juwara, Fida K. Dankar, Mark Clemons, Khaled El Emam

University of Ottawa · Children's Hospital of Eastern Ontario Research Institute · Charité - Universitaetsmedizin Berlin · McMaster University · The Ottawa Hospital Research Institute · Université de Montréal · Queen's University · University of Edinburgh · Ontario Institute for Cancer Research · University of Toronto · Leiden University Medical Center · St. Augustinus Hospital · University Center Mainz · National and Kapodistrian University of Athens

cs.LG, cs.CY

Submitted: 2025-09-27

Updated: 2026-08-18

Journal ref: J Med Internet Res 2026;28:e88665

DOI: 0.2196/88665

Code: https://github.com/pengpclab/PREDICTv3

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

Key concepts

Transfer Learning
This technique involves taking a model already trained on a large dataset and adapting it by fine-tuning it on a smaller, more specific dataset. It allows building upon existing knowledge rather than starting the training process from scratch.
PREDICT v3
This is an existing, widely-used prognostic tool that estimates survival rates for breast cancer patients. The paper used this tool as a baseline model to test whether it could be improved and adapted for specific patient populations.
Ensemble Approach
This method combines multiple different models (like fine-tuned PREDICT and Random Survival Forest) to create a single, more robust prediction tool. It improves reliability by allowing the system to fall back on alternative models when one fails.
External Validation
This process involves testing a model trained on one dataset (e.g., MA.twenty-seven) using data from an entirely different population or trial (e.g., TEAM). It determines if the model's improvements generalize beyond its original training setting.

Terminology

Summary

Summary

This study evaluates the potential of transfer learning, de-novo machine learning (ML), and ensemble integration to improve five-year survival prognostication in early breast cancer, comparing these approaches against the pre-trained prognostic tool PREDICT v3.

Background and Objectives: The paper notes that prognostic information is essential for decision-making in breast cancer management and that recent trials have predominantly focused on genomic prognostication tools, even though clinicopathological prognostication is less costly and more widely accessible. The authors state that "Advances in machine learning (ML), transfer learning and ensemble integration now offer opportunities to build robust prognostication frameworks, particularly in contexts where missingness and model assumptions vary across cohorts." The study specifically investigates four research questions: (1) whether fine-tuning PREDICT v3 to the MA.27 dataset improves survival prediction performance compared to the pre-trained model alone; (2) how state-of-the-art ML models trained directly on MA.27 compare against the fine-tuned PREDICT v3; (3) whether an ensemble of fine-tuned pre-trained models and de-novo ML models adds benefit compared to either approach alone; and (4) whether these potential benefits generalize to comparable external cohorts.

Methods: Data from the MA.27 trial (NCT00066573) was used for model training, with external validation on data from the TEAM trial (NCT00279448, NCT00032136) and a SEER cohort. The MA.27 dataset included 7,563 postmenopausal women with early-stage hormone receptor-positive breast cancer. The outcome was defined as breast cancer-related death within a 5-year observation interval. The models evaluated included: PREDICT v3 (the pretrained survival model), f-PREDICT v3 (the pretrained model fine-tuned to MA.27 via parameter-based transfer learning), Random Survival Forests (RSF), Extreme Gradient Boosting (XGB), and an ensemble integrating f-PREDICT v3, RSF, and XGB. The Integrated Calibration Index (ICI) was used as the optimization goal during training, with internal and external validation assessed in terms of calibration (ICI) and discrimination (area under the receiver operating characteristic curve, AUC). Shapley Additive Explanations (SHAP) were used to explain model predictions. The MA.27 dataset was split into 60% training, 20% testing, and 20% validation, with all steps repeated across 10 independent runs.

Key Results: The authors report that "Transfer learning, de-novo RSF, and ensemble integration relevantly improved calibration in MA.27 over the pre-trained model (ICI reduced from 0.042 in PREDICT v3 to ≤0.007) while discrimination remained comparable (AUC increased from 0.738 in PREDICT v3 to 0.744-0.799)." Specifically, f-PREDICT v3 achieved an ICI of 0.005 and AUC of 0.799, RSF achieved an ICI of 0.003 and AUC of 0.744, XGB achieved an ICI of 0.040 and AUC of 0.783, and the ensemble achieved an ICI of 0.007 and AUC of 0.746. The authors note that Invalid PREDICT v3 predictions were observed in 23.8-25.8% of MA.27 individuals due to missing information. In contrast, ML models and ensemble integration could predict survival regardless of missing information. When stratified by whether PREDICT v3 returned valid predictions, RSF and the ensemble still presented with good calibration (ICI 0.014 and 0.015 respectively) in the subset with invalid predictions.

Model Explainability: The SHAP analysis revealed that patient age, nodal status, pathological grading and tumor size had consistently highest SHAP values, indicating their importance for survival prognostication. In contrast, treatment information such as chemotherapy, radiotherapy or trastuzumab were typically ranked less important across all models.

External Validation: The authors report mixed results: External validation in SEER, but not in TEAM, confirmed the benefits of transfer learning, RSF and ensemble integration in terms of calibration. In the SEER cohort, f-PREDICT v3 achieved an ICI of 0.010 and AUC of 0.825, RSF achieved an ICI of 0.020 and AUC of 0.753, XGB achieved an ICI of 0.037 and AUC of 0.759, and the ensemble achieved an ICI of 0.018 and AUC of 0.792, all outperforming PREDICT v3 (ICI 0.039, AUC 0.765). However, in the TEAM cohort, "the MA.27 optimized models performed worse on the TEAM dataset with changes in AUC up to 0.122 (XGB from 0.783 in internal evaluation to 0.661 in external validation via TEAM) and changes in ICI up to 0.074 (f-PREDICT from 0.005 in internal evaluation to 0.079 in the external validation via TEAM)." In TEAM, PREDICT v3 outperformed all alternative approaches with an ICI of 0.034 and AUC of 0.701.

Conclusions: The authors conclude that this study demonstrates that transfer learning, de-novo RSF, and ensemble integration can improve prognostication in situations where relevant information for PREDICT v3 is lacking or where a dataset shift is likely. They further state that Ultimately, better survival estimation can provide meaningful guidance in breast cancer treatment, supporting a more targeted, cost-effective, and personalized approach to breast cancer care. The authors also note that the ensemble appeared to be the best approach given its ability to handle missingness, its superior performance in both calibration and discrimination compared to PREDICT-v3 and its comparable performance to f-PREDICT v3 in the SEER external validation context.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with what the improved system can do:


  • Improvement: Replace single-model approaches (like PREDICT v3) that fail when key variables are missing with an ensemble that dynamically falls back to models capable of handling missingness (RSF, XGB).

  • What the improved system can do: Provide valid 5-year survival predictions for 100% of patients, even when 23.8–25.8% of records lack critical information (e.g., tumor grade, nodal status). This eliminates the invalid prediction dead-end seen in PREDICT v3.

  • Improvement: Implement parameter-based transfer learning (fine-tuning) of pre-trained prognostic models (e.g., PREDICT v3) using a target cohort’s data. Optimize the 26 model parameters via gradient-free methods (Nelder-Mead) to minimize Integrated Calibration Index (ICI).

  • What the improved system can do: Reduce calibration error from ICI 0.042 (pre-trained) to 0.005 (fine-tuned) on MA.27, and from 0.039 to 0.010 on SEER—without sacrificing discrimination (AUC improved from 0.738 to 0.799). This yields clinically trustworthy probability estimates for decision-making.

  • Improvement: Train Random Survival Forests (RSF) and XGBoost with survival-specific loss functions, using tree-based methods that natively handle categorical variables, mixed missingness, and censored data. Optimize hyperparameters via grid search and use ICI as the training objective.

  • What the improved system can do: Achieve calibration (ICI 0.003 for RSF) and discrimination (AUC 0.744–0.799) comparable to or better than the fine-tuned PREDICT v3, while being fully applicable to patients with incomplete records. This is especially valuable for retrospective analyses and resource-limited settings.

  • Improvement: Combine f-PREDICT v3, RSF, and XGB predictions using convex weights optimized via Bayesian optimization. When f-PREDICT v3 cannot produce a prediction (due to missing inputs), the ensemble automatically re-weights to use only RSF and XGB.

  • What the improved system can do: Deliver a single, robust survival probability for every patient, with calibration (ICI 0.007) and discrimination (AUC 0.746) that are consistently near the best individual model—while never failing due to missing data. This is ideal for deployment in clinical decision support tools.

  • Improvement: Integrate model-agnostic Shapley Additive Explanations (SHAP) to quantify each variable’s contribution to individual predictions across all model types.

  • What the improved system can do: Automatically identify and rank the most influential prognostic factors (e.g., age, nodal status, tumor grade, tumor size) and show the direction of their effect (e.g., higher grade reduces survival probability). This enables clinicians to understand why a prediction was made, fostering trust and supporting shared decision-making.

  • Improvement: Build the system with a built-in external validation pipeline that tests the trained models on independent cohorts (e.g., SEER, TEAM) and reports both calibration (ICI) and discrimination (AUC) with bootstrapped confidence intervals.

  • What the improved system can do: Automatically flag when a model is likely to underperform due to dataset shift (e.g., TEAM showed AUC drop from 0.799 to 0.707 for f-PREDICT v3). This prevents deployment of overfitted models and guides when retraining or recalibration is necessary.

  • Improvement: Allow the system to be trained with either ICI (for decision-making tools) or AUC (for ranking/diagnostic tools) as the optimization goal, and expose the trade-off to the user.

  • What the improved system can do: Provide a model that is either well-calibrated (e.g., ICI 0.003 for RSF) for treatment planning, or highly discriminative (e.g., AUC 0.828 for f-PREDICT v3 with AUC optimization) for patient stratification—depending on the clinical use case.

  • Improvement: Implement automatic variable mapping and assumption-based construction (e.g., infer HER2 status from trastuzumab use, assume non-smoker status, approximate nodal count from TNM stage) to align any input dataset with the requirements of pre-trained models like PREDICT v3.

  • What the improved system can do: Accept raw clinical data (e.g., TNM stage, treatment history) and automatically transform it into the format required by the prognostic model, reducing manual preprocessing and enabling seamless integration with electronic health records.

  • Improvement: Include an optional ROSE (Random Over-Sampling Examples) module, but with a warning system that detects when re-balancing harms calibration (as seen in the paper: ICI increased from 0.003 to 0.247 for RSF).

  • What the improved system can do: Automatically recommend against re-balancing when the outcome is rare (e.g., 2.5% events) and the model is tree-based, preventing a common but harmful practice that degrades probability estimates.

  • Improvement: Use inverse probability of censoring weighting (IPCW) for AUC and hazard-regression-based smoothing for ICI, ensuring metrics are unbiased under heavy right-censoring (e.g., 97.5% censoring in MA.27).

  • What the improved system can do: Provide accurate performance evaluations that reflect true model capability, even in datasets where most patients do not experience the event within the follow-up window—critical for long-term survival studies.

Summary of Capabilities of the Improved AI System:

The improved system can predict 5-year breast cancer survival with high calibration (ICI ≤ 0.007) and discrimination (AUC ≥ 0.744) for all patients, including those with missing data. It can fine-tune pre-trained models to new cohorts, explain predictions via SHAP, validate on external data, and automatically choose the best training objective—making it a reliable, transparent, and generalizable tool for clinical decision support and retrospective research.

Abstract

Prognostic information is essential for decision-making in breast cancer management. Recently trials have predominantly focused on genomic prognostication tools, even though clinicopathological prognostication is less costly and more widely accessible. Machine learning (ML), transfer learning and ensemble integration offer opportunities to build robust prognostication frameworks. We evaluate this potential to improve survival prognostication in breast cancer by comparing de-novo ML, transfer learning from a pre-trained prognostic tool and ensemble integration. Data from the MA.27 trial was used for model training, with external validation on the TEAM trial and a SEER cohort. Transfer learning was applied by fine-tuning the pre-trained prognostic tool PREDICT v3, de-novo ML included Random Survival Forests and Extreme Gradient Boosting, and ensemble integration was realized through a weighted sum of model predictions. Transfer learning, de-novo RSF, and ensemble integration improved calibration in MA.27 over the pre-trained model (ICI reduced from 0.042 in PREDICT v3 to <=0.007) while discrimination remained comparable (AUC increased from 0.738 in PREDICT v3 to 0.744-0.799). Invalid PREDICT v3 predictions were observed in 23.8-25.8% of MA.27 individuals due to missing information. In contrast, ML models and ensemble integration could predict survival regardless of missing information. Across all models, patient age, nodal status, pathological grading and tumor size had the highest SHAP values, indicating their importance for survival prognostication. External validation in SEER, but not in TEAM, confirmed the benefits of transfer learning, RSF and ensemble integration. This study demonstrates that transfer learning, de-novo RSF, and ensemble integration can improve prognostication in situations where relevant information for PREDICT v3 is lacking or where a dataset shift is likely.

Related papers