D-JEPA: A Decision-Aligned Latent World Model
cs.RO, cs.LG
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 26 pages, including references and appendices. Project website: https://nebulis-lab.com/D-JEPA
Project page: https://nebulis-lab.com/D-JEPA
License: http://creativecommons.org/licenses/by/4.0/
The gist: Latent world models predict the consequences of actions, but accurate prediction does not guarantee that latent distance reflects which candidate will execute successfully.
Terminology
Abstract
Latent world models predict the consequences of actions, but accurate prediction does not guarantee that latent distance reflects which candidate will execute successfully. We identify a decision-local prediction gap: among the few futures competing for execution, a candidate predicted closer to the goal can produce a worse realized outcome than an available alternative. We introduce D-JEPA, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes. A bounded, permutation-equivariant operator jointly reasons over goal-relative predictive features and ordinal evidence, refining pretrained predictive geometry where action choices are most consequential. Restricted predictor adaptation and a shared ordinal interface extend this alignment across complementary predictive geometries. D-JEPA further realizes the learned decision structure in JEPA-compatible future representations, enabling deployment through native latent-distance planning. Evaluations across latent control, manipulation, pretrained action-producing models, physical robots and autonomous driving demonstrate improved action selection, including 87.89% success on PushT, a 15.04-point average gain on RoboTwin, and a 17-point gain on physical robot tasks. These results establish decision-relevant relational structure as a direct bridge between predictive world modeling and effective control.
Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control
- Fast LeWorldModel
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- DeepMind Control Suite
- What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?
- A Control Theory of Predictability in Latent World Models
- DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving