Clin-JEPA: A Multi-Phase Co-Training Framework for Joint-Embedding Predictive Pretraining on EHR Patient Trajectories
cs.LG, cs.AI, q-bio.QM
Submitted: 2026-05-11
Updated: 2026-09-16
Comments: 42 pages, 7 figures, 18 tables. Code: https://github.com/Kamaleswaran-Lab/Clin-JEPA
Code: https://github.com/Kamaleswaran-Lab/Clin-JEPA
License: http://creativecommons.org/licenses/by/4.0/
The gist: Joint-embedding predictive architectures (JEPA) learn representations by predicting in latent space, as in computer vision; retaining the action-conditioned predictor at inference turns them into
Terminology
Abstract
Joint-embedding predictive architectures (JEPA) learn representations by predicting in latent space, as in computer vision; retaining the action-conditioned predictor at inference turns them into latent world models, enabling planning in robotics (V-JEPA 2-AC). Bringing this design to EHR patient trajectories---a predictor that simulates a patient's trajectory in latent space---has not been explored. We use an LLM as the encoder, reading the hourly record as text, avoiding feature engineering and vocabulary harmonisation. But an LLM adapted by supervised fine-tuning does not organise its latent space around physiological dynamics, and freezing it to train the predictor, as in V-JEPA 2-AC, leaves the encoder unaware of the rollout signal: the predictor degrades under rollout. We instead co-train encoder and predictor under one latent-prediction objective, grounding the encoder in the dynamics its predictor must follow. Naïve co-training, however, is unstable: the untrained predictor drags the encoder toward collapse, and the predictor's rollout diverges as its target space moves. We present Clin-JEPA, a five-phase curriculum that stably co-trains an LLM encoder with a latent trajectory predictor on MIMIC-IV. Three evaluations support the design: (1) under 48-hour autoregressive rollout the co-trained predictor degrades least (predictor degradation times 1.06, against times 1.23--1.36 for two-stage designs and times 6.3--66 for curriculum ablations) while the co-trained encoder resolves the progression of patient state most sharply (largest state displacement); (2) the co-trained encoder separates deteriorating from stable patients in its latent space with Cohen's d = 1.59, against at most 0.50 for two-stage encoders; (3) one set of embeddings serves 34 downstream tasks across three benchmarks, outperforming strong per-task tuned baselines and a pretrained EHR foundation model.
Sources
- EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis
- Revisiting Feature Prediction for Learning Visual Representations from Video
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- EHRMamba: Towards Generalizable and Scalable Foundation Models for Electronic Health Records
- Building the EHR Foundation Model via Next Event Prediction
- Foundation models for electronic health records: representation dynamics and transferability
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- The Patient is not a Moving Document: A World Model Training Paradigm for Longitudinal EHR
- Mastering Atari with Discrete World Models
- CLARITY: Medical World Model for Guiding Treatment Decisions by Simulating Context-Aware Disease Trajectories
- Beyond Generative AI: World Models for Clinical Prediction, Counterfactuals, and Planning
- An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks