Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in Cardiology
Yunsung Chung, Yingshuo Liu, Abboud F. Hassan, Han Feng, Mary M. Maleckar, Nassir Marrouche, Jihun Hamm
Tulane University · Simula Research Laboratory
cs.LG, cs.CV
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: Medical World Models (MWM) Workshop at MICCAI 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper proposes an intervention-aware clinical world model for forecasting post-operative outcomes in cardiology, specifically applied to atrial fibrillation (AF) ablation.
Terminology
Summary
The paper proposes an intervention-aware clinical world model for forecasting post-operative outcomes in cardiology, specifically applied to atrial fibrillation (AF) ablation. The model represents each patient with a structured 3D latent state derived from pre-ablation LGE-MRI, then evolves this state through time-ordered post-intervention events during the 90-day blanking period. The model updates the latent state using procedural context (ablation heatmap), static covariates, elapsed time, and peri-event physiological embeddings from ECGFounder. A follow-up MRI provides training-only supervision through a latent forecasting objective, matching the predicted terminal latent state to the encoded follow-up MRI latent. The model forecasts 451-day AF recurrence and scar extent using only observations available by a selected query horizon, without requiring follow-up MRI intensities at inference.
The method uses a frozen variational autoencoder pretrained on a larger imaging cohort (N=732) to map MRIs to 3D spatial latent states. The initial state is set to the encoded pre-ablation MRI, and the training-only target is the encoded follow-up MRI. Events are tokenized with normalized event time, normalized queried horizon, binary flags for medication, cardioversion, or repeat procedure, medication category embeddings, and mean-pooled ECGFounder embeddings from within 7 days before each event. A Transformer encoder contextualizes the event sequence, and a 3D residual CNN with a context-dependent MLP drift updates the latent state per token, with elapsed-time drift scaled by a 30-day time constant. The terminal horizon token advances the state without adding a clinical event, allowing queries at different horizons by changing only the time in this token.
On the complete-record cohort (N=91, 5-fold cross-validation with 3 seeds), the model achieves AUROC 0.756 ± 0.051 and AUPRC 0.777 ± 0.058 for recurrence prediction, improving on matched-input LSTM by 0.103/0.090. The model also achieves a scar-extent MAE of 2.971 ± 0.675 percentage points, slightly better than the oracle-style Post-MRI-only baseline (3.189 ± 0.698) despite not receiving follow-up MRI intensities at inference. The OOF Brier score is 0.201 and five-bin expected calibration error is 0.032. Anytime risk trajectories show that separation develops as blanking-period evidence accumulates from D30 to D90, with risk rising fastest for early recurrence, remaining stable for no recurrence, and intermediate for late recurrence.
Ablation studies reveal a contribution hierarchy: removing events or latent matching produces the largest AUROC reductions (0.133 and 0.132), while removing the ablation map or ECG has smaller effects. Without latent matching, AUROC SD rises from 0.051 to 0.200, indicating structural supervision stabilizes the learned state. Direct fusion loses 0.047 AUROC and 0.057 AUPRC, supporting step-wise evolution over one-shot concatenation. Shuffling event content while preserving times reduces AUROC/AUPRC from 0.756/0.777 to 0.709/0.713, indicating event identity contributes beyond elapsed time. Input-editing sensitivity tests show heterogeneous model dependence on blanking-period scenarios, with a right-skewed distribution of risk differences (mean 0.136) driven by a subset of higher-risk patients.
On an auxiliary cohort without usable ablation geometry (N=258), the model variant without an ablation map achieves AUROC/AUPRC 0.713/0.747, gains of 0.158/0.155 over the strongest matched-input sequence baseline, resembling the N=91 no-map ablation (0.711/0.758). Limitations include the small complete-record cohort (N=91), internal evaluation only within DECAAF-II, potential manual EAM-to-MRI mapping errors, and the associational rather than causal nature of input edits. The paper concludes that the framework supports multi-horizon risk updates and retrospective input edits, but larger full-modality cohorts and external validation are needed before extending toward causal modeling.
Improvements for AI systems
Improvements to AI Systems:
- Add a latent-state forecasting objective with structural supervision.
Train the model to predict a terminal latent state that matches an encoded follow-up image, rather than only optimizing the final clinical label. This stabilizes learned representations (AUROC SD drops from 0.200 to 0.051) and improves generalization.
- Implement step-wise temporal evolution of latent states via a residual 3D CNN with context-dependent drift.
Replace one-shot concatenation of all inputs with a sequential update that processes each post-intervention event in time order. This yields +0.047 AUROC and +0.057 AUPRC over direct fusion.
- Use a frozen pretrained variational autoencoder to map imaging to structured 3D latent states.
Pretrain on a larger cohort (N=732) and freeze the encoder. This provides a consistent, high-dimensional spatial representation that supports both initial-state encoding and training-only target supervision, without needing follow-up images at inference.
- Tokenize clinical events with rich, time-aware embeddings.
Encode each event using normalized event time, queried horizon, medication/cardioversion/repeat-procedure flags, medication category embeddings, and mean-pooled physiological embeddings from a foundation model (e.g., ECGFounder). This improves event identity beyond elapsed time alone (shuffling content drops AUROC by 0.047).
- Introduce a terminal horizon token for multi-horizon queries.
Add a non-clinical token that advances the latent state to a queried time point. This allows the same model to generate anytime risk trajectories (e.g., D30, D60, D90) by changing only the time in that token, enabling dynamic risk updates without retraining.
- Enable retrospective input-editing sensitivity analysis.
Because the model evolves latent states from events, you can edit or remove specific inputs (e.g., ablation heatmap, ECG embeddings) and re-run inference to quantify their causal influence on risk. This provides heterogeneous, patient-specific dependence profiles (e.g., right-skewed risk differences) for interpretability.
- Add a training-only latent matching loss to reduce variance and improve calibration.
Use the follow-up MRI latent as a supervision signal only during training. This reduces overfitting, improves expected calibration error (0.032), and yields a Brier score of 0.201, making the model more reliable for clinical decision support.
What the Improved AI System Can Do:
-
Predict post-operative AF recurrence and scar extent at multiple time horizons (e.g., 90-day blanking period, 451-day follow-up) using only pre-ablation MRI and event logs, without requiring follow-up MRI at inference.
-
Generate anytime risk trajectories that show how risk evolves as clinical evidence accumulates, enabling early identification of high-risk patients (e.g., fastest risk rise for early recurrence).
-
Provide stable, interpretable latent representations that are robust to small cohort sizes (N=91) and support transfer to auxiliary cohorts lacking certain modalities (e.g., no ablation map) with minimal performance loss (AUROC 0.713 vs. 0.711).
-
Perform counterfactual-style
what-if
analyses by editing input events (e.g., removing a cardioversion or changing medication timing) to assess the impact on predicted outcomes, aiding personalized treatment planning. -
Achieve state-of-the-art calibration and discrimination (AUROC 0.756, AUPRC 0.777) while outperforming sequence baselines (LSTM) by 0.103/0.090, and even slightly beat an oracle that uses post-MRI intensities for scar extent prediction (MAE 2.971 vs. 3.189).
Abstract
Many clinical prediction models treat post-intervention outcomes as a one-step mapping from baseline measurements to a future endpoint. However, recovery after a procedure often unfolds as an irregular trajectory: clinical observations, medication changes, repeat interventions, and physiological measurements are recorded asynchronously and can change risk assessment over time. We propose an intervention-aware clinical world model that represents each patient with a structured latent state and evolves it through time-ordered post-intervention events. The model first encodes baseline imaging into a 3D spatial latent state. It then updates this state using procedural context, static covariates, elapsed time, and peri-event physiological embeddings. Follow-up imaging provides training-only supervision through a latent forecasting objective. We apply the framework to atrial fibrillation ablation. During the 90-day recovery window, irregular post-procedure records provide clinically meaningful evidence for long-term recurrence risk. In repeated internal cross-validation on DECAAF-II, our model achieves AUROC 0.756 and AUPRC 0.777 for recurrence prediction. It also achieves a scar-extent MAE of 2.971 percentage points without requiring follow-up MRI intensities at inference. The learned state supports recurrence-risk queries at different horizons and retrospective input editing of blanking-period records.
Sources
- Revisiting Feature Prediction for Learning Visual Representations from Video
- MuDreamer: Learning Predictive World Models without Reconstruction
- CRAFT: Clinical Reward-Aligned Finetuning for Medical Image Synthesis
- CLARITY: Medical World Model for Guiding Treatment Decisions by Simulating Context-Aware Disease Trajectories
- An Electrocardiogram Foundation Model Built on over 10 Million Recordings with External Evaluation across Multiple Domains
- EHRWorld: A Patient-Centric Medical World Model for Long-Horizon Clinical Trajectories
- Beyond Generative AI: World Models for Clinical Prediction, Counterfactuals, and Planning
- Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning
- Generative Artificial Intelligence in Medical Imaging: Foundations, Progress, and Clinical Translation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks