V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
cs.CV, cs.AI, cs.RO
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/breez3young/VJEPA-Policy
Terminology
Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Revisiting Feature Prediction for Learning Visual Representations from Video
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- Causal World Modeling for Robot Control
- Video Generators are Robot Policies
- JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling
- Flow Matching Guide and Code
- LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
- VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
- Wan: Open and Advanced Large-Scale Video Generative Models
- Qwen-Image Technical Report
- InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning
- ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models