CoJEPA: Combining Contrastive Learning and JEPA for Global-Local Music Representations
cs.SD, cs.AI, cs.LG, eess.AS, eess.SP
Submitted: 2026-08-31
Updated: 2026-09-06
Terminology
Sources
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
- Contrastive Learning of Musical Representations
- Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation
- MATPAC++: Enhanced Masked Latent Prediction for Self-Supervised Audio Representation Learning
- Understanding self-supervised Learning Dynamics without Contrastive Pairs
- Toward Fully Self-Supervised Multi-Pitch Estimation
- W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training
- Feature Normalization Prevents Collapse of Non-contrastive Learning Dynamics
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- The Llama 3 Herd of Models
- Parameter-Efficient Transfer Learning for Music Foundation Models
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment