FOCAL: Fine-Grained Optimal-Transport-Driven Contrastive Alignment of Language and ECGs with Waveform Enhancement
eess.SP, cs.AI, cs.LG
Submitted: 2025-05-17
Updated: 2026-09-17
Comments: EMNLP 2026
Code: https://github.com/meta-llama/llama3
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Electrocardiograms (ECGs) are essential non-invasive tools for diagnosing cardiovascular diseases.
Terminology
Abstract
Electrocardiograms (ECGs) are essential non-invasive tools for diagnosing cardiovascular diseases. While recent multimodal ECG-Report contrastive learning methods have shown promise for zero-shot ECG interpretation, they predominantly rely on global representations, failing to capture the fine-grained relationship between localized waveform patches and specific pathological tags. This limitation is further exacerbated by the fact that nearly 55% of standard clinical reports (e.g., in MIMIC-ECG) lack explicit waveform descriptions. In this paper, we propose FOCAL, a novel framework that achieves precise, fine-grained alignment between localized ECG segments and individual report tags via Optimal Transport. Furthermore, because fine-grained alignment at the tag level exacerbates the false negative problem among reports sharing common diagnoses, we introduce a semantic similarity matrix to guide the contrastive objective and correct misalignments. To address the scarcity of detailed waveform text, we introduce a coarse-to-fine enrichment pipeline that leverages Large Language Models (LLMs) to recover missing semantics, utilizing a coarse model verification step to rigorously filter out hallucinations. Extensive experiments across six datasets demonstrate that FOCAL establishes new state-of-the-art performance in zero-shot prediction and linear probing.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Publicly Available Clinical BERT Embeddings
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Time-Series Representation Learning via Temporal and Contextual Contrasting
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Mistral 7B
- TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
- FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment
- BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains
- Modeling Caption Diversity in Contrastive Vision-Language Pretraining
- Zero-Shot ECG Classification with Multimodal Learning and Test-time Clinical Knowledge Enhancement
- Guiding Masked Representation Learning to Capture Spatio-Temporal Relationship of Electrocardiogram
- Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey
- Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding
- MedCLIP: Contrastive Learning from Unpaired Medical Images and Text
- ECG Semantic Integrator (ESI): A Foundation ECG Model Pretrained with LLM-Enhanced Cardiological Text
- Self-Supervised Time Series Representation Learning via Cross Reconstruction Transformer
Related papers
- Runtime Assurance Under Measurement Attack: Necessary and Sufficient Observability Conditions for Learned Control in Radio Access Networks
- Physics-Constrained Deep Learning Model for Contactless Blood Pressure Monitoring from Triaxial Bodyseismography
- Uncertainty Quantification in Machine Learning for Biosignal Applications -- A Review
- Continuous Orthogonal Mode Decomposition: Haptic Signal Prediction in Tactile Internet
- Generative Models for Modeling and Synthesizing MIMO Channels in Adverse Weather Conditions
- Deep-Learning-Based Pixelated Microwave Filter Design and Characterization using Electro-Optical Electric-Field Measurements