US-JEPA: A Joint Embedding Predictive Architecture for Ultrasound
cs.CV, cs.AI, cs.LG
Submitted: 2026-02-22
Updated: 2026-09-07
Code: https://github.com/ButterflyNetwork/MITGrandHack2018
License: http://creativecommons.org/licenses/by/4.0/
The gist: Ultrasound (US) imaging poses unique challenges for representation learning due to its inherently noisy acquisition process.
Terminology
Abstract
Ultrasound (US) imaging poses unique challenges for representation learning due to its inherently noisy acquisition process. The low signal-to-noise ratio and stochastic speckle patterns hinder standard self-supervised learning methods relying on a pixel-level reconstruction objective. Joint-Embedding Predictive Architectures (JEPAs) address this drawback by predicting masked latent representations rather than raw pixels. However, standard approaches depend on hyperparameter-brittle and computationally expensive online teachers updated via exponential moving average. We propose US-JEPA, a self-supervised framework that adopts the Static-teacher Asymmetric Latent Training (SALT) objective. By using a frozen, domain-specific teacher to provide stable latent targets, US-JEPA decouples student-teacher optimization and pushes the student to expand upon the semantic priors of the teacher. In addition, we provide the first rigorous comparison of all publicly available state-of-the-art ultrasound foundation models on UltraBench, a public dataset benchmark spanning multiple organs and pathological conditions. Under linear probing for diverse classification tasks, US-JEPA achieves performance competitive with or superior to domain-specific and universal vision foundation model baselines. Our results demonstrate that masked latent prediction provides a stable and efficient path toward robust ultrasound representations.
Sources
- Understanding intermediate layers using linear classifier probes
- General Methods Make Great Domain-specific Foundation Models: A Case-study on Fetal Ultrasound
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- Emerging Properties in Self-Supervised Vision Transformers
- A Simple Framework for Contrastive Learning of Visual Representations
- Momentum Contrast for Unsupervised Visual Representation Learning
- Masked Image Modeling: A Survey
- Segment Anything
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
- A New Dataset and A Baseline Model for Breast Lesion Detection in Ultrasound Videos
- Tumor Detection, Segmentation and Classification Challenge on Automated 3D Breast Ultrasound: The TDSC-ABUS Challenge
- USF-MAE: Ultrasound Self-Supervised Foundation Model with Masked Autoencoding
- DINOv2: Learning Robust Visual Features without Supervision
- DINOv3
- The Open Kidney Ultrasound Data Set
- GraphEcho: Graph-Driven Unsupervised Domain Adaptation for Echocardiogram Video Segmentation
- A Fully Open and Generalizable Foundation Model for Ultrasound Clinical Applications
- MMOTU: A Multi-Modality Ovarian Tumor Ultrasound Image Dataset for Unsupervised Cross-Domain Semantic Segmentation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models