Fine-tuning an ECG Foundation Model to Predict Coronary CT Angiography Outcomes

arXiv:2512.05136 · cs.CV, cs.AI · Submitted 2026-08-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Fine-tuning an ECG Foundation Model to Predict Coronary CT Angiography Outcomes".

Jane: , utilizing only information contained within the text:

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, looking at the title, "Fine-tuning an ECG Foundation Model to Predict Coronary CT Angiography Outcomes," it really tells us exactly what they’re doing: taking a large existing model designed for one thing and adapting it for a more specific, useful purpose. The authors are quite a big team from Peking University and Tianjin Medical University.

Jane: That title is very descriptive; it highlights the core challenge they tackled, which is using an ECG to predict what CCTA shows about vessel blockages. It makes the purpose of the research immediately clear for anyone listening.

Lu: The authors are clearly drawing on a lot of expertise across different domains, which I think is a good sign for tackling such a complex problem; it shows they aren't just looking at one narrow angle.

Meng: I wonder how much effort went into fine-tuning that foundation model; adapting it to predict specific vessel outcomes sounds like it requires very precise data handling and validation processes.

Lalam: It’s impressive seeing such a collaborative effort between institutions; the breadth of expertise involved suggests they built a very robust system, not just a quick proof of concept.

The paper's summary: Tom: To summarize what this paper is doing, they used consecutive multicenter clinical data from four distinct sources to build and test their AI-ECG model. They aimed to predict hemodynamically significant stenosis in specific arteries, defined by certain percentage thresholds like at least seventy percent stenosis in the Right Coronary Artery or Left Anterior Descending artery.

Jane: That’s a lot of data input for one model; they built this AI-ECG model using a transfer learning framework to fine-tune an ECG foundation model, which lets it make vessel-specific predictions instead of just looking at the patient as a whole.

Lu: The methodology is strong because they didn't just train it on one dataset; they used internal pairs from Peking University, external validation cases from Tianjin Medical University, and a longitudinal cohort of four hundred patients. That multi-center design is crucial for showing generalizability across different patient populations.

Meng: I see the use of those specific thresholds—seventy percent in RCA or LAD, and fifty percent in the Left Main artery—that gives the prediction a very clear clinical target to aim for; it grounds the abstract model into tangible medical goals.

Lalam: The paper also showed they used calibration analysis and decision curve analysis to see how well their AI-derived risk strata aligned with existing clinical risk assessments, which is really important for showing real-world utility.

The paper's improvements: Tom: One of the key parts of this work is how they moved beyond just a single prediction and designed a fusion strategy that integrated the AI model’s risk strata with established guideline-based Pre-Test Probability categories. This combination actually improved the Negative Predictive Value and achieved a positive Net Reclassification Improvement compared to using the guideline categories alone.

Jane: That fusion technique is clever because it doesn't discard established clinical knowledge; it combines what the AI found with what clinicians already know, making the final risk stratification much more reliable for patient care.

Lu: They also conducted waveform-based and attribution-based analyses to characterize ECG morphology differences between groups and identify which specific signal regions contributed most strongly to those model predictions. This is a vital step because it gives us a window into the mechanism behind the prediction, moving it past just being a black box.

Meng: From an engineering standpoint, knowing *why* the model made a prediction by looking at those attribution maps helps us build more trustworthy systems; it moves the AI from being just an output generator to something we can actually debug and trust when things go wrong.

Lalam: I think this focus on interpretability is what really elevates this research; when clinicians can see the specific ECG segments driving a high-risk score, it builds the necessary trust for them to adopt these new tools in their daily practice.

Conclusion: Tom: So, to wrap up the paper "Fine-tuning an ECG Foundation Model to Predict Coronary CT Angiography Outcomes," the main implication is that this AI-ECG approach offers a feasible tool for complementary CAD screening and anatomical risk estimation, especially in resource-constrained settings. They showed that their fusion strategy provides greater net clinical benefit than using any single method alone.

Jane: It really shows how we can take existing diagnostic tools and augment them with AI to get more detailed, vessel-specific information from simple signals like ECGs, which is a significant step toward better personalized treatment plans.

Lu: The longitudinal follow-up data was also very telling; the Kaplan-Meier analysis clearly showed that the high-risk group consistently had the highest event rate over eight hundred days, which confirms that their risk stratification actually predicts long-term adverse cardiovascular events.

Meng: The paper does point out a limitation, though; they note that definitive diagnosis still requires further integration into routine clinical practice, meaning this is clearly meant to be a supportive tool rather than the final word on diagnosis right now.

Lalam: Overall, "Fine-tuning an ECG Foundation Model to Predict Coronary CT Angiography Outcomes" offers a low-barrier, high-impact pathway for improving patient triage and risk assessment using existing clinical data streams.

Tom: That's all we have time for today; we’ve seen how this work on fine-tuning the AI model can make a real difference in identifying heart disease risks. We'll be back next time with more exciting research from arXiv.

Yujie Xiao, Qinghao Zhao, Gongzheng Tang, Hao Zhang, Zhuoran Kan, Deyun Zhang, Jun Li, Guangkun Nie, Xiaocheng Fang, Haoyu Wang, Shun Huang, Tong Liu, Jian Liu*, Kangyin Chen*, Shenda Hong*

Institute of Medical Technology and Peking University Health Science Center · National Institute of Health Data Science and Peking University · Department of Cardiology at Peking University People’s Hospital · Tianjin Key Laboratory of Ionic-Molecular Function of Cardiovascular Disease, Department of Cardiology, Tianjin Institute of Cardiology, The Second Hospital of Tianjin Medical University · Heart Voice Medical Technology · School of Intelligence Science and Technology at Peking University · University of Chinese Academy of Sciences · State Key Laboratory of Vascular Homeostasis and Remodeling and National Health Center Key Laboratory of Cardiovascular Molecular Biology and Regulatory Peptides at Peking University · Institute for Artificial Intelligence at Peking University

cs.CV, cs.AI

Submitted: 2026-08-21

Updated: 2026-08-24

Code: https://github.com/shunxio/AnyECG-CCTA

Importance score: 79/100

The gist: " While coronary computed tomographic angiography (CCTA) is a primary diagnostic tool, its widespread application is limited by "resource requirements and radiation exposure," as well as the risk of

Key concepts

ECG Foundation Model Fine-tuning
This involves adapting a large existing model trained on electrocardiogram data to make specific predictions about coronary CT angiography outcomes. The researchers used transfer learning to fine-tune this foundation model for vessel-specific predictions instead of general patient assessments.
Fusion Strategy
This technique combines the AI model's risk strata with established guideline-based Pre-Test Probability categories. This combination was shown to improve the Negative Predictive Value and achieve a positive Net Reclassification Improvement compared to using guideline categories alone.
Interpretability (Attribution-based Analyses)
This involves using waveform-based and attribution-based analyses to characterize ECG morphology differences between groups. This helps identify which specific signal regions in the ECG contributed most strongly to the model's predictions, making the AI less of a 'black box' and more trustworthy for clinicians.
Longitudinal Follow-up Data
The study used longitudinal follow-up data, including a Kaplan-Meier analysis. This analysis confirmed that patients in the high-risk group consistently had the highest event rate over eight hundred days, validating that their risk stratification predicts long-term adverse cardiovascular events.

Terminology

Summary

The following is a detailed summary of the scientific paper, utilizing only information contained within the text:

Context and Problem Statement

Coronary atherosclerotic heart disease (CAD) represents a major global public health burden. While coronary computed tomographic angiography (CCTA) is a primary diagnostic tool, its widespread application is limited by resource requirements and radiation exposure, as well as the risk of contrast-induced nephropathy. This necessitates the development of complementary, low-cost, and non-invasive methods for CAD risk assessment.

Study Objective and Methodology

The study aimed to develop and validate an AI-ECG model using CCTA as an anatomical reference standard to predict vessel-specific coronary stenosis. The model was designed to identify hemodynamically significant stenosis defined as:

  • at least 70% stenosis in the Right Coronary Artery (RCA), Left Anterior Descending (LAD), and Left Circumflex (LCX.

  • at least 50% stenosis in the Left Main artery (LM).

The researchers utilized a transfer learning framework to fine-tune an ECG foundation model, ECGFounder, which had been adapted for MI-related prediction tasks. This framework allowed the model to provide vessel-specific predictions rather than only patient-level assessment.

Data Characteristics and Cohort Design

The study employed a multi-center design involving three distinct cohorts:

  1. Internal Dataset: 4,620 internal ECG-CCTA pairs from Peking University People’s Hospital (n=4,620).

  2. External Validation Dataset: 2,477 external validation cases from the Second Hospital of Tianjin Medical University (n=2,477).

  3. Longitudinal Follow-up Cohort: 400 patients from the Second Hospital of Tianjin Medical University (n=400).

The study's design ensured that the distribution of multi-label coronary artery outcomes remained balanced across the validation folds.

Model Performance and Discrimination

The AI-ECG model demonstrated stable discriminative performance:

  • In internal validation, the model achieved a micro-averaged area under the receiver operating characteristic curve (micro-AUC) of 0.732 for patient-level hemodynamically significant CAD. Individual vessel AUC values were 0.744 (RCA), 0.683 (LM), 0.716 (LAD), and 0.736 (LCX, respectively).

  • In the external validation set, the model achieved an AUC of 0.694 for detecting patient-level significant CAD, with corresponding vessel-specific AUC values of 0.714 (RCA), 0.740 (LAD), 0.700 (LCX), and 0.673 (LM).

The model maintained stable performance across various subgroups, including age, sex, and the interval between ECG and CCTA acquisition.

Performance in Clinically Normal ECG Subgroups

A key finding was the model's ability to detect risk even when standard interpretation fails. The model maintained an AUC of 0.710 for identifying patient-level significant CAD in the subgroup of clinically normal ECG interpretations, demonstrating its potential value for screening occult CAD.

** Clinical Risk Stratification and Utility**

The continuous model-predicted probabilities were converted into three clinically interpretable risk strata: low, intermediate, and high risk. This was achieved using two distinct thresholds:

  1. Low-risk threshold (t low): Selected under a sensitivity constraint of 90% to support safe rule-out decisions.

  2. High-risk threshold (t high):): Selected under a specificity constraint of 95% to support reliable identification.

  • Calibration and Decision Curve Analysis (DCA): Calibration curves showed agreement between predicted and observed risk. DCA indicated that the model-based risk stratification provided greater net clinical benefit compared to both treat-all and treat-none strategies.

  • Fusion Strategy: The study integrated the AI model's risk strata with guideline-based Pre-Test Probability (PTP) categories (low: at most 5%; intermediate: 5%–15%; high: at least 15%). This fusion strategy improved Negative Predictive Value (NPV) and achieved a positive Net Reclassification Improvement (NRI) compared to using PTP alone.

** Longitudinal Prognostic Assessment**

In the longitudinal follow-up cohort, Kaplan-Meier analysis showed a clear separation of major adverse cardiovascular event (MACE) risk across the three model-defined risk groups over 800 days. The high-risk group consistently exhibited the highest event rate, while the low-risk group maintained a very low cumulative event incidence.

** Model Interpretability and Mechanism**

To understand how the model predicts risk, researchers performed waveform and attribution analyses:

  • Waveform Analysis: The high-risk group showed the largest deviation from the low-risk template, indicating that the predicted risk ordering is reflected in a graded change of waveform morphology rather than an abrupt binary transition.

  • Attribution Analysis: Using Integrated Gradients (IG), the attribution maps showed that salient contributions were temporally concentrated around structured ECG intervals, with the densest attribution bands appearing near the QRS complex and its adjacent repolarization-related regions, confirming that the model is anchored to physiologically meaningful waveform segments.

Conclusion and Future Outlook

The findings support AI-ECG as a feasible tool for complementary CAD screening, anatomical risk estimation, and clinical triage. The study concludes that the fusion strategy allows for a low-barrier, high-impact clinical utility in resource-constrained settings. However, the authors note that definitive diagnosis still requires further integration into routine clinical practice and emphasized that a prospective randomized controlled trial is needed to confirm its clinical impact.

Improvements for AI systems

The core weakness in current AI systems, as evidenced by these references, is often a siloed approach—treating each diagnostic task (e.g., CAD detection, MI mapping) independently. To mitigate the risk of false negatives and to achieve true clinical utility, the system must evolve from a Diagnostic Tool to an Integrated Predictive Clinical Decision Support System.

Here are three critical improvements that must be implemented in the AI architecture:


Improvement: Instead of training separate models for ECG analysis, heart sound analysis, and structural diagnosis, we must build a single Foundation Model (building upon concepts from Citations 1065 and 1018). This model must be trained simultaneously on diverse modalities (ECG signals, phonocardiogram recordings, potentially retinal images/labs) to learn shared underlying physiological representations.

Technical Specificity:

  • The architecture must incorporate a Transformer-based encoder stack that processes time-series data from multiple sources (e.g., 12-lead ECG and heart sound spectrograms) in parallel before fusing them at a central bottleneck layer.

  • We must implement Multi-Task Learning (MTL), leveraging techniques like those described in Citations 1075 and 1072, where the model optimizes for multiple related outcomes simultaneously (e.g., predicting CAD risk and estimating kidney function and detecting heart failure severity).

What the Improved System Can Do:

  • Holistic Risk Profiling: It can generate a single, comprehensive Composite Risk Score that weighs evidence from all available modalities. For example, it won't just say CAD present, but rather: Moderate to High risk of future MI (Composite Score: 0.85), driven by combination of T-wave inversion (ECG) and reduced S3 gallop sound (Phono).

  • Causal Feature Extraction: By utilizing techniques like Gradient Surgery (Citation 1075), the model can be forced to highlight the specific, most physiologically relevant features in the raw data that drove a high-risk prediction, significantly improving interpretability for clinicians.

Sources

Related papers