Latent evolving World Action Model
cs.CV, cs.RO
Submitted: 2026-09-23
Updated: 2026-09-24
Code: https://github.com/XuejiFang/LeWAM
Terminology
Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- RT-H: Action Hierarchies Using Language
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RT-1: Robotics Transformer for Real-World Control at Scale
- What Makes Pre-Trained Visual Representations Successful for Robust Manipulation?
- LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
- DINOv3-Diffusion Policy: Self-Supervised Large Visual Model for Visuomotor Diffusion Policy Learning
- Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning
- Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
- GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
- SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
- OpenVLA: An Open-Source Vision-Language-Action Model
- Depth Anything 3: Recovering the Visual Space from Any Views
- DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- JEPA-VLA: Video Predictive Embedding is Needed for VLA Models
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- Octo: An Open-Source Generalist Robot Policy
- Intelligent Switching for Reset-Free RL
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models