Transfer Learning and Machine Learning for Training Five Year Survival Prognostic Models in Early Breast Cancer

summary

Video file (mp4)

In short

The episode discusses a paper using transfer learning and machine learning to improve five-year survival prediction for early breast cancer. Hosts review how fine-tuning existing models can boost accuracy, discuss the benefits of ensemble approaches for handling missing data, and emphasize that external validation is crucial because models are not universally applicable.

Key concepts

Transfer Learning
This technique involves taking a model already trained on a large dataset and adapting it by fine-tuning it on a smaller, more specific dataset. It allows building upon existing knowledge rather than starting the training process from scratch.
PREDICT v3
This is an existing, widely-used prognostic tool that estimates survival rates for breast cancer patients. The paper used this tool as a baseline model to test whether it could be improved and adapted for specific patient populations.
Ensemble Approach
This method combines multiple different models (like fine-tuned PREDICT and Random Survival Forest) to create a single, more robust prediction tool. It improves reliability by allowing the system to fall back on alternative models when one fails.
External Validation
This process involves testing a model trained on one dataset (e.g., MA.twenty-seven) using data from an entirely different population or trial (e.g., TEAM). It determines if the model's improvements generalize beyond its original training setting.

Terminology used across episodes

This episode discusses

The paper

Transfer Learning and Machine Learning for Training Five Year Survival Prognostic Models in Early Breast Cancer · Read on arXiv

Lisa Pilgram, Kai Yang, Ana-Alicia Beltran-Bless, Gregory R. Pond, Lisa Vandermeer, John Hilton, Marie-France Savard, Andréanne Leblanc, Lois Sheperd, Bingshu E. Chen, John M. S. Bartlett, Karen J. Taylor, Jane Bayani, Sarah L. Barker, Melanie Spears, Cornelis J. H. van der Velde, Elma Meershoek-Klein Kranenbarg, Luc Dirix, Elizabeth Mallon, Annette Hasenburg, Christos Markopoulos, Lamin Juwara, Fida K. Dankar, Mark Clemons, Khaled El Emam

University of Ottawa · Children's Hospital of Eastern Ontario Research Institute · Charité - Universitaetsmedizin Berlin · McMaster University · The Ottawa Hospital Research Institute · Université de Montréal · Queen's University · University of Edinburgh · Ontario Institute for Cancer Research · University of Toronto · Leiden University Medical Center · St. Augustinus Hospital · University Center Mainz · National and Kapodistrian University of Athens

Prognostic information is essential for decision-making in breast cancer management. Recently trials have predominantly focused on genomic prognostication tools, even though clinicopathological prognostication is less costly and more widely accessible. Machine learning (ML), transfer learning and ensemble integration offer opportunities to build robust prognostication frameworks. We evaluate this potential to improve survival prognostication in breast cancer by comparing de-novo ML, transfer learning from a pre-trained prognostic tool and ensemble integration. Data from the MA.27 trial was used for model training, with external validation on the TEAM trial and a SEER cohort. Transfer learning was applied by fine-tuning the pre-trained prognostic tool PREDICT v3, de-novo ML included Random Survival Forests and Extreme Gradient Boosting, and ensemble integration was realized through a weighted sum of model predictions. Transfer learning, de-novo RSF, and ensemble integration improved calibration in MA.27 over the pre-trained model (ICI reduced from 0.042 in PREDICT v3 to <=0.007) while discrimination remained comparable (AUC increased from 0.738 in PREDICT v3 to 0.744-0.799). Invalid PREDICT v3 predictions were observed in 23.8-25.8% of MA.27 individuals due to missing information. In contrast, ML models and ensemble integration could predict survival regardless of missing information. Across all models, patient age, nodal status, pathological grading and tumor size had the highest SHAP values, indicating their importance for survival prognostication. External validation in SEER, but not in TEAM, confirmed the benefits of transfer learning, RSF and ensemble integration. This study demonstrates that transfer learning, de-novo RSF, and ensemble integration can improve prognostication in situations where relevant information for PREDICT v3 is lacking or where a dataset shift is likely.

DOI: 0.2196/88665

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Transfer Learning and Machine Learning for Training Five Year Survival Prognostic Models in Early Breast Cancer".

Jane: The paper was written by Lisa Pilgram, Kai Yang, Ana-Alicia Beltran-Bless, Gregory R. Pond, Lisa Vandermeer et al. from University of Ottawa and Children's Hospital of Eastern Ontario Research Institute and Charité - Universitaetsmedizin Berlin and McMaster University and The Ottawa Hospital Research Institute and Université de Montréal and Queen's University and University of Edinburgh and Ontario Institute for Cancer Research and University of Toronto and Leiden University Medical Center and St. Augustinus Hospital and University Center Mainz and National and Kapodistrian University of Athens.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. We've got a fascinating paper on the table today, and it's all about breast cancer survival prediction. The title is "Transfer Learning and Machine Learning for Training Five Year Survival Prognostic Models in Early Breast Cancer." Jane, I have to say, when I first saw this title, I thought, okay, this is a technical methods paper, but the more I dig in, the more I realize how much this could actually change clinical practice.

Jane: Absolutely, Tom. And for our listeners who might not be deep in the weeds of machine learning, let's break down what that title actually means. "Transfer learning" is essentially taking a model that's already been trained on a huge amount of data, and then fine-tuning it on a smaller, more specific dataset. Think of it like learning to drive a car, and then learning to drive a truck — you don't start from scratch, you build on what you already know.

Tom: Right, and that's exactly what they did here. They took an existing, widely-used prognostic tool called PREDICT v3, which estimates survival for breast cancer patients, and they fine-tuned it on data from a clinical trial called MA.twenty-seven. The idea being, can we make this general tool better by adapting it to a specific patient population?

Jane: And that's the key question, because these pre-trained models are built on one population, but the patients you see in your clinic might be different. The authors are essentially asking, can we bridge that gap? And the answer, based on what we're seeing, is a resounding yes, at least in some cases.

Tom: Yeah, the results are pretty striking. The fine-tuned model, they call it f-PREDICT v3, saw its calibration error drop from zero point zero four two down to zero point zero zero five. For the non-experts, that means the model's predicted survival probabilities got much, much closer to what actually happened to the patients.

Jane: And that's not just a statistical nicety. If you're a doctor telling a patient, "you have an eighty percent chance of surviving five years," you want that number to be accurate. A miscalibrated model could lead to over-treatment or under-treatment, which is a huge deal.

Tom: Exactly. And what's interesting is that they didn't just stop at fine-tuning. They also trained completely new models from scratch on the MA.twenty-seven data, using something called Random Survival Forests and XGBoost, which are pretty standard machine learning tools. And they compared all of these against the original, un-tuned PREDICT v3.

Jane: So we've got a three-way comparison happening here: the old model, the fine-tuned model, and the brand-new models. And the results, as we'll get into in a bit, are not as simple as "one size fits all." There are trade-offs, especially when you start looking at external validation on other datasets.

Tom: Right, and that's where it gets really interesting, because a model that works great on one population might fall apart on another. We'll get into that in the next segment, but for now, let's just say this paper is asking the right questions about how we build and adapt these tools.

Jane: And it's a question that matters for so many diseases, not just breast cancer. If we can figure out how to adapt these models effectively, we can make personalized medicine more accessible and more accurate for everyone.

Tom: Stay tuned, because we're about to dig into the actual methods and results. You don't want to miss this.

Summary: Tom: Welcome back. We're diving deeper into "Transfer Learning and Machine Learning for Training Five Year Survival Prognostic Models in Early Breast Cancer." Last segment, we talked about the core idea of fine-tuning a pre-trained model. Now, let's get into what the paper actually did and found.

Jane: Right, and the first thing that stands out is the data. They used the MA.twenty-seven trial, which had over seven thousand five hundred postmenopausal women with early-stage, hormone receptor-positive breast cancer. That's a very specific group, and it's a great test case because it's a real clinical trial with real outcomes.

Tom: And the outcome they were predicting is five-year survival, which is a key milestone in breast cancer. But here's the thing — the dataset was heavily imbalanced. Only one hundred eighty-seven patients, about two point five percent, actually had a breast cancer-related death within that five-year window. So the model has to learn from a very small number of events.

Jane: That's a classic problem in survival analysis, and it's one of the reasons why they considered re-balancing the data. They tested a technique called ROSE, which essentially creates synthetic examples of the minority class to balance things out. But interestingly, it didn't help. In fact, it made calibration much worse.

Tom: Yeah, that was a surprising negative result. The calibration error went from something like zero point zero zero five up to zero point three or higher in some models. So they wisely dropped that approach and just trained on the original, imbalanced data.

Jane: And that's a good lesson for anyone doing this kind of work — sometimes the standard tricks don't work, and you have to trust the data as it is. But the bigger story here is the comparison between the models. The fine-tuned PREDICT v3 and the Random Survival Forest both performed really well in terms of calibration, with ICI values around zero point zero zero three to zero point zero zero five.

Tom: For our listeners, that ICI is the Integrated Calibration Index, and lower is better. A value of zero point zero zero five means the average difference between predicted and observed survival is half a percent. That's incredibly accurate.

Jane: And discrimination, which is the model's ability to tell who will survive and who won't, was also solid. The AUC values ranged from about zero point seven four to zero point eight zero. The fine-tuned model actually had the best discrimination at zero point seven nine nine, which is a meaningful improvement over the original PREDICT v3 at zero point seven three eight.

Tom: But here's where it gets complicated. They also did external validation on two other datasets — SEER, which is a big US cancer registry, and TEAM, which is another clinical trial. And the results were mixed. On SEER, the fine-tuned model and the others still looked good. On TEAM, they didn't.

Jane: That's the reality check. The MA.twenty-seven-optimized models actually performed worse on TEAM than the original PREDICT v3. The AUC dropped, and the calibration got worse. So the improvements they saw in MA.twenty-seven didn't fully generalize.

Tom: And that's a crucial finding, because it tells us that transfer learning isn't a magic bullet. It works when the new population is similar to the training population, but when there's a big shift, it can actually hurt. We'll talk more about what that means for practice in the next segment.

Jane: For now, the takeaway is that this paper is a rigorous, honest look at what works and what doesn't when you're trying to build these prognostic models. It's not just a success story — it's a nuanced picture of the challenges.

Improvements: Tom: Welcome back to our discussion of "Transfer Learning and Machine Learning for Training Five Year Survival Prognostic Models in Early Breast Cancer." We've covered the basics and the mixed results on external validation. Now, let's talk about what this paper actually suggests we should do differently.

Jane: Right, and one of the most practical improvements is handling missing data. The original PREDICT v3 simply couldn't generate a prediction for about a quarter of the MA.twenty-seven patients — between twenty-three point eight percent and twenty-five point eight percent — because key information like tumor grade or size was missing.

Tom: That's a huge gap. If you're a clinician and the tool just says "sorry, I can't help you" for one in four patients, that's a problem. But the machine learning models, like the Random Survival Forest and the ensemble, could handle missing data internally. They could predict survival for every single patient.

Jane: And that's a real improvement. The tree-based models use something called surrogate splits, which essentially find alternative ways to split the data when a variable is missing. So they don't just throw up their hands — they work with what they have.

Tom: And the ensemble, which combines the fine-tuned PREDICT, the Random Survival Forest, and XGBoost, was designed to fall back on the machine learning models when PREDICT couldn't make a prediction. So you get the best of both worlds — the accuracy of the fine-tuned model when it works, and the robustness of the ML models when it doesn't.

Jane: That's a really smart design. And the paper also looked at which variables matter most for survival prediction. Using SHAP analysis, they found that patient age, nodal status, tumor grade, and tumor size were consistently the most important factors across all models.

Tom: That makes sense clinically. Those are the classic prognostic factors that oncologists have been using for decades. But it's nice to see that the machine learning models are picking up on the same signals, which gives us confidence that they're learning something real.

Jane: And interestingly, treatment variables like chemotherapy and radiotherapy were ranked lower in importance. That doesn't mean they don't matter, but it suggests that for predicting overall survival, the tumor biology and patient characteristics dominate.

Tom: Now, the paper also suggests that if you're going to use these models in a new population, you need to be careful. The external validation on TEAM showed that the MA.twenty-seven-tuned models didn't transfer well. So the improvement here is really about knowing when to trust your model and when to re-tune it.

Jane: And that's a practical takeaway for anyone building these tools. You can't just train once and deploy everywhere. You need to validate on your own population, and if it doesn't work, you need to fine-tune again.

Tom: Exactly. And the authors are pretty clear that this isn't a one-size-fits-all solution. But the framework they've laid out — fine-tuning, ensemble integration, and careful external validation — is a solid blueprint for building better prognostic models.

Jane: And that's the real contribution here. It's not just about breast cancer. It's about how we should approach model development and adaptation in medicine more broadly.

Conclusion: Tom: And that brings us to the end of our discussion on "Transfer Learning and Machine Learning for Training Five Year Survival Prognostic Models in Early Breast Cancer." Jane, what's the big picture here?

Jane: The big picture, Tom, is that this paper gives us a realistic and practical roadmap for improving survival prediction in breast cancer. Fine-tuning an existing tool like PREDICT v3 can dramatically improve its accuracy on a new population, and machine learning models can fill in the gaps when data is missing.

Tom: And the ensemble approach — combining the fine-tuned model with the ML models — gives you a robust tool that can handle real-world messiness. But the external validation results remind us that these models aren't universal. They need to be tested and adapted for each new setting.

Jane: Right. The paper is honest about the limitations. The TEAM dataset didn't show the same benefits, which tells us that dataset shift is a real challenge. But the framework they've built — train, fine-tune, ensemble, validate — is exactly what we need to move forward.

Tom: And for patients, this could mean more accurate survival estimates, which can guide treatment decisions and reduce both over-treatment and under-treatment. That's a meaningful impact.

Jane: Absolutely. And I think the biggest takeaway is that we have the tools to make these models better, but we have to be thoughtful about how we use them. It's not just about throwing more data at a model — it's about adapting and validating.

Tom: Well said, Jane. We've covered a lot of ground today, from the basics of transfer learning to the nuances of external validation. I hope our listeners found this as fascinating as we did.

Jane: And we're already looking forward to the next paper. Thanks for tuning in, everyone. We'll see you next time.

Tom: Take care, and keep questioning the numbers.

More episodes

← Home