CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation".
Jane: The paper was written by the authors from Eindhoven University of Technology and Singapore Management University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Jane: Following up on that idea of creating a shared language for cardiac data, let's talk about what the paper summarizes regarding this shared representation. They use a technique called JEPA, which seems to be central to how they achieve this invariance.
Tom: Right, so we're moving past just understanding the title and looking at the methodology itself. The core finding highlighted in their shared space analysis is that cross-modal training makes the pooled cardiac code invariant to sensing modality, dropping the silhouette from zero point one two one down to-zero point zero zero six—that's a massive number change!
Lu: That transition from zero point one two one down to negative suggests a profound level of alignment. It implies that even when the measurements are fundamentally different—say, one is voltage-based and another is light absorption based—the resulting shared code treats them as describing the same underlying physical process.
Meng: That invariance figure is really interesting from an engineering standpoint because it quantifies robustness. If the model doesn't care *which* sensor generated the input, only that it's cardiac related, that dramatically reduces our dependency on perfect sensor calibration across different hospital sites.
Lalam: From a cultural perspective, this resilience means that healthcare infrastructure can finally achieve global standardization in remote monitoring without worrying about proprietary hardware compatibility breaking the core diagnostic logic.
Jane: So, to simplify what they're showing with this shared space: it means the model has figured out the universal features of a heartbeat that exist *despite* the differences between PPG and ECG inputs, right?
Tom: Exactly! And I want to circle back to Lu’s point about invariance. It’s not just that it *can* process both; it's that it processes them so they become indistinguishable in the shared latent space, which is what makes the code robust.
Lu: Because the mathematical proof of invariance suggests that the model has successfully stripped away all modality-specific noise and artifacts, leaving only the true signal structure common to all inputs.
Meng: If we're thinking about implementation, does this shared representation also help with dealing with missing data? If one sensor drops out entirely, can the model still predict something useful based on the others because of this universal code?
Lalam: I believe the implication here is that it elevates AI in medicine from being merely descriptive—just telling you what happened—to being predictive and unifying, creating a cohesive digital twin of cardiac health across multiple data streams.
Improvements: Tom: We’ve seen the theoretical elegance of this shared space, so now we're looking at the practical improvements shown by comparing Stage I and Stage II checkpoints. The paper claims this comparison isolates the benefit purely from cross-modal pretraining, which is smart methodology.
Jane: That isolated contribution is key, because it lets us say definitively: "Look, *this* pretraining step made a measurable difference." They show this with two very different tasks: atrial fibrillation and the six-class arrhythmia task.
Lu: Let's focus on the binary AFib task first; seeing the class silhouette climb from zero point zero three at Stage I to zero point one two at Stage II is a tangible visualization of improved separation power, which is fantastic work.
Meng: Quantitatively, that jump in the class silhouette suggests that the boundaries between 'normal' and 'AFib' are becoming much cleaner in the feature space after Stage II training. Are we talking about improving sensitivity or specificity most significantly here?
Lalam: The improvement in class structure signals a move toward higher diagnostic confidence. For patients
Paper discussion segment 3: Tom: So, we’ve seen how effective the shared space is under ideal lab conditions using CardioState-JEPA, but what about when those perfect conditions aren't met in a real hospital?
Jane: Exactly. The biggest leap forward isn't just that it works; it's how robustly it handles failure or interference from different sensors.
Lu: That’s where the true power of cross-modal learning shines, I think; if one modality gets corrupted by noise, the model shouldn't collapse entirely because the other signals are carrying redundant, complementary information.
Meng: You hit on a critical point there regarding noise; practically speaking, we need to know what happens when a patient is moving or when the ECG electrodes aren't making perfect contact—the system has to degrade gracefully.
Lalam: Graceful degradation implies resilience, which means the AI needs to interpret *intent* from incomplete data, not just patterns from complete data.
Tom: Right, so we’re moving from "perfect beat detection" to "best possible cardiac estimate even with missing beats," which is a huge operational jump.
Jane: And that requires the delay aligner—the whole point of CardioState-JEPA—to be flexible enough to assume an approximate timing rather than requiring a precise, golden reference point every single time.
Lu: I wonder if we could extend this concept to model *drift* over long periods, not just beat-to-beat alignment; maybe tracking subtle changes in the physiological offset itself?
Meng: If you’re modeling drift, are you talking about the delay changing because the patient's underlying condition is changing—say, a developing arrhythmia—or just sensor shift?
Jane: It could be both, but thinking about clinical utility, addressing sensor shift is probably more immediately impactful for getting this into bedside monitors.
Tom: So essentially, we’re building a meta-model that not only understands the heart's rhythm but also understands the *machine's* limitations and how to compensate for them.
Lalam: That ability to self-correct based on input quality is fundamentally improving human-computer interaction in medicine, allowing clinicians to trust the AI even when they feel uneasy about the raw signal quality.
Lu: We could design a confidence score layer that weights the reliability of each modality's contribution dynamically, telling the doctor exactly how much they can trust the prediction right now.
Meng: That confidence score is crucial; if it just spits out a number without saying, "Warning: ECG data compromised," it’s useless because doctors won't know where to look for errors.
Jane: So the output isn't just a classification, but a full risk assessment package that details *why* it thinks what it thinks.
Tom: That makes perfect sense; we’re not just diagnosing the heart; we’re diagnosing the data stream itself and how it relates to the heart.
Lalam: This shift towards transparent uncertainty quantification means AI won't be a black box, but an assistant that helps build trust by acknowledging its own boundaries.
Meng: Okay, knowing all this robustness talk, what’s next? Are we talking about integrating this into wearable patches or dedicated bedside units first?
Conclusion: Tom: So, wrapping up our deep dive on cross-modal cardiac representations, it really strikes you just how much these models are pushing the boundaries of what we thought was possible with multi-sensor data.
Jane: Exactly, Tom. What's so impressive is that they didn't just make the signals align; they learned *why* and *how* they align—the true electromechanical timing between ECG, PPG, and PCG.
Lu: I mean, the idea that you can take multiple streams of data—electrical signals, light pulses, sound waves—and fuse them into one shared cardiac code that’s immune to which sensor you used is revolutionary for diagnostics.
Meng: But Lu's point about the practical implementation sticks out to me; if the model can achieve that invariant representation, it means we could build devices that don't even need perfect calibration between different hardware types.
Lalam: And building on Meng’s thought, this kind of robust, multi-modal AI framework has massive implications beyond just diagnosing arrhythmias; it fundamentally changes how remote patient monitoring works by creating a single source of cardiac truth.
Jane: That's such a powerful way to put it, Lalam; it elevates the entire field from data collection to true predictive understanding.
Tom: It really feels like we’ve seen the blueprint for the next generation of wearable health tech with this paper, "CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation."
Meng: I agree with Tom; the robustness they demonstrated, especially against noise and sensor variability, is what moves this from an academic curiosity to a deployable product line.
Lu: Thinking about the future work they mentioned, imagine applying this architecture not just to routine monitoring but to real-time intervention guidance during surgery.
Lalam: And when we consider the cultural impact—how much anxiety or uncertainty this could remove for millions of people who currently rely on intermittent, fragmented diagnoses—it’s genuinely profound.
Tom: It’s certainly exciting stuff, Jane; I think we've got a lot to digest from this work before we sign off today.
Jane: Thanks so much to all of you for walking us through the intricacies of this research!
Eindhoven University of Technology · Singapore Management University
cs.LG, eess.IV, stat.ML
Submitted: 2026-08-13
Updated: 2026-09-09
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: CardioState-JEPA is a cardiac foundation model that learns a single shared representation jointly across electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG), built on a
Key concepts
- Cross-Modal Learning
- This technique allows the model to learn a single, shared understanding of cardiac data by processing multiple types of inputs simultaneously. It ensures that the resulting code treats different measurements, like voltage-based or light absorption based, as describing the same underlying physical process.
- Shared Cardiac Representation
- This is a universal digital code or latent space that captures the core features of a heartbeat regardless of which sensor generated the data. The goal is to create a single source of 'cardiac truth' immune to variations in hardware or input type.
- Invariance
- In this context, invariance means that the model's output remains stable and consistent even when the input data changes fundamentally (e.g., switching from ECG to PPG). This quantifies robustness by stripping away modality-specific noise and artifacts.
Terminology
Summary
CardioState-JEPA is a cardiac foundation model that learns a single shared representation jointly across electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG), built on a physiology-aware joint-embedding predictive architecture. The model maps heterogeneous waveforms into a common token space, processes them with a single shared Transformer encoder, and learns by predicting masked latent cardiac states, placing the pretraining target on shared physiology rather than sensor-specific waveform appearance. To handle the temporal offsets between electrical, mechanical, and hemodynamic events, cross-modal prediction uses a learned delay aligner that matches signals at the corresponding cardiac time. Because synchronized multi-sensor recordings are scarce, CardioState-JEPA first learns within-modality structure from abundant unimodal data and then uses paired data to align modalities in latent cardiac time.
The paper states: "We introduce CardioState-JEPA, a cardiac foundation model to learn a single shared representation jointly across ECG, PPG, and PCG, built on a physiology-aware joint-embedding predictive architecture. The model maps heterogeneous waveforms into a common token space, processes them with a single shared Transformer encoder, and learns by predicting masked latent cardiac states, placing the pretraining target on shared physiology rather than sensor-specific waveform appearance."
The model is motivated by the observation that ECG, PPG, and PCG provide complementary views of the same cardiac cycle, yet existing cardiac foundation models are trained for a single sensing modality, leaving the shared physiology across sensors unexploited.
The paper notes that "Electrical activation initiates contraction, valve motion then produces the heart sounds, and blood ejection appears later as a peripheral pulse, so the three modalities observe the same beat at different physiological times."
The architecture consists of modality-specific tokenizers that map each signal into a common token space, followed by a single shared Transformer encoder. The paper explains: "Given xm ∈ R Cm ×Tm, the tokenizer produces hm = fm (xm) ∈ R N ×d: a strided convolution whose stride is set per modality so all three emit tokens at a comparable rate, a multi-scale depthwise block that adds context around short events such as the QRS complex, PPG upstroke, and first heart sound, and a linear projection, followed by sinusoidal positional and modality embeddings. A single shared Transformer encoder then processes these tokens with no modality-specific layers, so ECG, PPG, and PCG update the same parameters."
The training uses a two-stage curriculum. Stage I learns within-modality structure from abundant unimodal data through intra-modal masked latent prediction. Stage II uses paired recordings to align modalities through delay-aware cross-modal prediction. The paper describes: Stage I learns each modality's structure from abundant unimodal data, and Stage II aligns modalities from scarce paired data through a learned delay.
For the intra-modal objective, the paper states: "We predict masked regions in latent space rather than reconstructing the waveform, because reconstruction rewards fitting sensor noise and fine morphology, exactly the nuisance we want the code to forget. Following the joint-embedding predictive principle, we mask a large fraction of the tokens in contiguous blocks, encode the visible context, and let a light predictor q theta match the momentum encoder's code at the masked positions."
For the cross-modal objective, the paper explains: "Aligning tokens at equal timestamps would match an electrical event to a mechanical or hemodynamic one and teach surface appearance rather than shared state, so we predict through a learned delay. A delay head h delta maps the source code zim and a target-modality embedding en to a bounded per-token offset. The target is gathered with a Gaussian kernel centered at the shifted time, and the delay is supervised with physiological anchors:
beat detection gives a reference offset per pair, the R-peak to first-heart-sound interval for ECG–PCG and the pulse arrival time for ECG–PPG."
The full objective combines several terms: L = lambdaintra Lintra + lambdacross Lcross + lambdadelay Ldelay-sup + lambdastate Lstate + lambdaphase Lphase.
The paper notes: In Stage I, each sample is a single modality, so only intra-modal prediction and the phase term are active; the cross-modal, delay, and state terms require paired inputs and switch on in Stage II.
The pretraining data includes MIMIC-IV-ECG with 800K 12-lead ECG recordings at 500 Hz, PPG-EXT as a large-scale PPG corpus at 125 Hz, and BMD-HS for PCG recordings at 4000 Hz. For cross-modal supervision, the paper uses PPG-EXT and VitalDB for synchronous ECG-PPG pairs, EPHNOGRAM for ECG-PCG pairs, and SensSmartTech for trimodal recordings.
The model is evaluated as a frozen encoder across 25 downstream tasks spanning ECG, PPG, and PCG. The paper reports: "our encoder improves average PPG classification by 8.2 AUROC points, PCG murmur detection by 18.8 AUROC points, and ECG classification by 15.5 AUROC points over the best self-supervised signal baseline and matches or exceeds cardiac models trained with privileged clinical text or supervised labels on several ECG benchmarks."
Specifically, on PPG tasks, CardioState-JEPA improves the average classification AUROC from 72.2 to 80.4 compared with the strongest baseline, while also reducing the average regression MAE from 10.9 to 9.1.
For ECG, CardioState-JEPA improves the average AUROC across all 18 settings by 15.5 points over the strongest self-supervised baseline (MoCo-v3; 84.1 vs. 68.5).
For PCG, the model achieves the best results on both CirCor murmur detection and CinC2016 abnormal heart-sound detection under the linear probing protocol.
The paper also visualizes the shared space using t-SNE: "after Stage I the samples form clear modality-specific clusters, since each modality has only been trained in isolation. After Stage II the same samples become thoroughly interleaved across modalities, indicating that cross-modal training reshapes the space from a sensor-driven layout into a modality-invariant one. The modality silhouette
falls from 0.121 after Stage I to −0.006 after Stage II."
Ablation studies show that the full trimodal setting gives the best overall balance, masked latent prediction gives the strongest results compared to SimCLR, BYOL, BarlowTwins, and MAE, and that cross-modal prediction, state alignment, delay modeling, and phase supervision provide complementary benefits.
The paper concludes: "We introduced CardioState-JEPA, a physiology-aware joint embedding predictive framework that learns from ECG, PPG, and PCG through intra-modal latent prediction and delay-aware cross-modal alignment. By predicting latent cardiac states and aligning modalities at the corresponding cardiac time, CardioState-JEPA enables electrical, hemodynamic, and acoustic signals to supervise one another. Our experiments show that the unified model transfers across modalities and surpasses strong self-supervised signal baselines on representative tasks."
Improvements for AI systems
Improvements to AI Systems:
- Cross-Modal Latent State Prediction with Delay-Aware Alignment
-
Implement a joint-embedding predictive architecture that maps heterogeneous sensor inputs (e.g., ECG, PPG, PCG) into a shared token space, then predicts masked latent cardiac states rather than raw waveforms.
-
Use a learned delay aligner to match tokens at corresponding physiological times (e.g., R-peak to pulse arrival), avoiding naive timestamp alignment that would mix electrical, mechanical, and hemodynamic events.
-
Resulting capability: An AI system that can fuse asynchronous multi-sensor data (e.g., wearable ECG + PPG + microphone) into a single physiology-aware representation, enabling robust cardiac monitoring even when sensors are not perfectly synchronized.
- Two-Stage Curriculum Learning for Scarce Paired Data
-
First pretrain on abundant unimodal data using intra-modal masked latent prediction (Stage I), then fine-tune on limited paired multi-modal data using delay-aware cross-modal prediction (Stage II).
-
This reduces reliance on expensive synchronized recordings while still achieving modality-invariant representations.
-
Resulting capability: An AI system that can be trained on large-scale single-modality datasets (e.g., millions of ECGs) and then adapted to new modalities (e.g., PPG or PCG) with only a small paired dataset, dramatically lowering data collection costs for multi-modal health monitoring.
- Physiology-Anchored Delay Supervision
-
Supervise cross-modal delay predictions using known physiological intervals (e.g., R-peak to first heart sound for ECG-PCG, pulse arrival time for ECG-PPG) derived from beat detection.
-
This grounds the model in cardiac physiology rather than arbitrary signal alignment.
-
Resulting capability: An AI system that can automatically estimate physiological timing parameters (e.g., pulse transit time, electromechanical delay) from raw waveforms, enabling non-invasive hemodynamic and cardiac function assessment without manual feature engineering.
- Modality-Invariant Representation via Shared Transformer Encoder
-
Use a single shared Transformer encoder with no modality-specific layers, after modality-specific tokenizers that normalize sampling rates and token densities.
-
This forces the model to learn a common cardiac state representation across sensors.
-
Resulting capability: An AI system that can perform zero-shot or few-shot transfer across modalities—e.g., a model trained on ECG can directly classify PPG or PCG signals, or a system can switch sensors (e.g., from chest patch to wristband) without retraining.
- Masked Latent Prediction to Ignore Sensor Noise
-
Replace waveform reconstruction with latent-space prediction of masked tokens, using a momentum encoder for target codes.
-
This discards sensor-specific noise and fine morphology, focusing on shared physiological state.
-
Resulting capability: An AI system that is robust to sensor artifacts, baseline wander, and noise, improving performance on real-world wearable data where signal quality varies.
- Multi-Scale Tokenization for Short Cardiac Events
-
Use strided convolutions with per-modality stride and multi-scale depthwise blocks to capture short events (QRS complex, PPG upstroke, S1 heart sound) while maintaining comparable token rates across modalities.
-
Resulting capability: An AI system that can accurately detect and classify brief, transient cardiac events (e.g., arrhythmias, murmurs) even when they occur at different temporal resolutions across sensors.
- Joint State and Phase Supervision
-
Incorporate auxiliary losses for cardiac state (e.g., beat phase) and phase alignment, in addition to intra- and cross-modal prediction.
-
This provides explicit physiological supervision during pretraining.
-
Resulting capability: An AI system that can segment cardiac cycles into phases (systole, diastole) and estimate cardiac timing intervals directly from raw signals, enabling automated echocardiography-like measurements from simple wearables.
- Frozen Encoder Transfer Across 25 Downstream Tasks
-
Evaluate the pretrained encoder without fine-tuning on diverse tasks (classification, regression) across ECG, PPG, and PCG.
-
This demonstrates that the learned representation is general-purpose and task-agnostic.
-
Resulting capability: An AI system that can be deployed as a universal cardiac feature extractor, enabling rapid development of new diagnostic tools (e.g., detecting hypertension, sleep apnea, or heart failure) by simply adding a linear classifier on top of the frozen encoder, without task-specific pretraining.
What the Improved AI System Can Do:
-
Monitor cardiac health continuously from a single wearable device that combines ECG, PPG, and acoustic sensors, fusing them into a unified physiological state representation.
-
Detect and classify a wide range of cardiac conditions (arrhythmias, murmurs, abnormal heart sounds) with higher accuracy than single-modality systems, even in noisy, real-world conditions.
-
Estimate physiological parameters (pulse transit time, electromechanical delay, cardiac phase) non-invasively and in real-time.
-
Transfer learned knowledge across sensor types and data distributions, reducing the need for large annotated multi-modal datasets.
-
Provide robust performance on unseen devices or recording conditions due to noise-invariant latent representations.
Abstract
Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the same cardiac cycle, yet existing cardiac foundation models are trained for a single sensing modality, leaving the shared physiology across sensors unexploited. We introduce CardioState-JEPA, a cardiac foundation model to learn a single shared representation jointly across ECG, PPG, and PCG, built on a physiology-aware joint-embedding predictive architecture. The model maps heterogeneous waveforms into a common token space, processes them with a single shared Transformer encoder, and learns by predicting masked latent cardiac states, placing the pretraining target on shared physiology rather than sensor-specific waveform appearance. To handle the temporal offsets between electrical, mechanical, and hemodynamic events, cross-modal prediction uses a learned delay aligner that matches signals at the corresponding cardiac time. Because synchronized multi-sensor recordings are scarce, CardioState-JEPA first learns within-modality structure from abundant unimodal data and then uses paired data to align modalities in latent cardiac time. Evaluated as a frozen encoder across 25 downstream tasks spanning ECG, PPG, and PCG, our encoder improves average PPG classification by 8.2 AUROC points, PCG murmur detection by 18.8 AUROC points, and ECG classification by 15.5 AUROC points over the best self-supervised signal baseline and matches or exceeds cardiac models trained with privileged clinical text or supervised labels on several ECG benchmarks. These results establish that heterogeneous cardiac signals can mutually supervise a single foundation model of cardiac physiology.
Sources
- MOMENT: A Family of Open Time-series Foundation Models
- Chronos-2: From Univariate to Universal Forecasting
- AnyPPG: An ECG-Guided PPG Foundation Model Trained on Over 100,000 Hours of Recordings for Holistic Health Profiling
- PulseLM: A Foundation Dataset and Benchmark for PPG-Text Learning
- StethoLM: Audio Language Model for Cardiopulmonary Analysis Across Clinical Tasks
- CoST: Contrastive Learning of Disentangled Seasonal-Trend Representations for Time Series Forecasting
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks