WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving
cs.CV, cs.AI
Submitted: 2026-08-21
Updated: 2026-09-05
Code: https://github.com/AFARI-Research/WA-JEPA
License: http://creativecommons.org/licenses/by/4.0/
The gist: Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent feature prediction.
Terminology
Abstract
Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent feature prediction. However, V-JEPA is built around random-mask completion and deterministic regression, making it fundamentally ill-suited for autonomous driving planning that demands future-directed prediction tightly coupled with action. To address this, we rethink the V-JEPA paradigm and present WA-JEPA, a V-JEPA-native world-action model designed for autonomous driving planning. Instead of random spatiotemporal masking, WA-JEPA employs hybrid future-masked pre-training, where the model infers future latents from observed context. Departing from deterministic regression, we recast future prediction as conditional flow matching over latent futures, which substantially improves the model's ability to generate plausible future latents for downstream planning. Finally, a joint future-action predictor is proposed to denoise future scene tokens and ego trajectories together in a unified spatiotemporal latent space, allowing action supervision to directly shape planning-relevant world representations. Pre-trained on nuPlan videos and fine-tuned on NAVSIM, WA-JEPA reaches 91.7 EPDMS on NAVSIM-v2, surpassing the strongest end-to-end and world-action baselines by 1.6 and 1.3 EPDMS, and, without HUGSIM-specific fine-tuning, attains the best HD-Score of 0.4462 on the closed-loop HUGSIM benchmark under the same evaluation protocol. These results validate V-JEPA-native world-action modeling as a powerful and scalable paradigm for autonomous driving planning. Code is available at https://github.com/AFARI-Research/WA-JEPA.
Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Revisiting Feature Prediction for Learning Visual Representations from Video
- Pseudo-Simulation for Autonomous Driving
- NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking
- ARTEMIS: Autoregressive End-to-End Trajectory Planning with Mixture of Experts for Autonomous Driving
- LTX-Video: Realtime Video Latent Diffusion
- DriveFuture: Future-Aware Latent World Models for Autonomous Driving
- CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving
- EMMA: End-to-End Multimodal Model for Autonomous Driving
- DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
- ImagiDrive: A Unified Imagination-and-Planning Framework for Autonomous Driving
- Hydra-MDP++: Advancing End-to-End Driving via Expert-Guided Hydra-Distillation
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
- ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
- DriveVA: Video Action Models are Zero-Shot Drivers
- DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving
- SparseDriveV2: Scoring is All You Need for End-to-End Autonomous Driving
- Wan: Open and Advanced Large-Scale Video Generative Models
- Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
- Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models