Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
cs.SD, cs.AI, cs.LG, eess.AS, eess.SP
Submitted: 2025-06-25
Updated: 2026-09-26
Code: https://github.com/LudovicTuncay/Audio-JEPA
Terminology
Sources
- HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
- CED: Consistent ensemble distillation for audio tagging
- Scaling up masked audio encoder learning for general audio classification
- Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
- Revisiting Feature Prediction for Learning Visual Representations from Video
- Masked Autoencoders that Listen
- TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems
- GeoJEPA: Towards Eliminating Augmentation- and Sampling Bias in Multimodal Geospatial Learning
- Decoupled Weight Decay Regularization
- Clotho: An Audio Captioning Dataset
- FMA: A Dataset For Music Analysis
- General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline
- FSD50K: An Open Dataset of Human-Labeled Sound Events
- Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders
- ConvFormer: Plug-and-Play CNN-Style Transformers for Improving Medical Image Segmentation
- MetaFormer Baselines for Vision
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment