Subspace-Decomposed JEPAs: Disentangling Progression and Content in Latent World Models
cs.LG
Submitted: 2026-05-29
Updated: 2026-09-15
Code: https://github.com/LucasStill/SD-JEPA
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Joint-Embedding Predictive Architectures (JEPAs) learn compact latent world models by predicting future embeddings, but no single coordinate of the latent is designated to encode task progression.
Terminology
Abstract
Joint-Embedding Predictive Architectures (JEPAs) learn compact latent world models by predicting future embeddings, but no single coordinate of the latent is designated to encode task progression. We carve the JEPA latent into two orthogonal subspaces with disjoint roles: a low-dimensional progression subspace shaped by a cosine-margin triplet loss, and a high-dimensional content subspace regularised by the existing SIGReg objective of LeWM. We prove that the two anti-collapse forces act on disjoint coordinates, so they compose additively rather than competing on the same dimensions. Our method, SD-JEPA improves over the LeWM baseline on the majority of its control benchmarks at matched compute, and outperforms the strongest non-LeWM JEPA baseline on Push-T; a subspace-ablation falsifier confirms the split is the load-bearing ingredient. Beyond planning, the resulting 1-D angular progression coordinate functions as a scene-aware compass on the latent. It advances with task progress, regresses when the agent backtracks, and under controlled perturbations both spikes and relocalises to a semantically appropriate new task-phase sector, separating the moment of surprise from its meaning in a way that prediction-error scalars cannot.
Sources
- Cosmos World Foundation Model Platform for Physical AI
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
- Learning Latent Action World Models In The Wild
- Mastering Diverse Domains through World Models
- AI-Generated Video Detection via Perceptual Straightening
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- OGBench: Benchmarking Offline Goal-Conditioned RL
- GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
- Joint Embedding Predictive Architectures Focus on Slow Features
- DeepMind Control Suite
- HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels
- Temporal Straightening for Latent Planning
- Latent Action Pretraining from Videos
- DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks