From World Models to World Action Models: Rethinking Next-State Prediction
cs.RO
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA
- World Models
- Dream to Control: Learning Behaviors by Latent Imagination
- Mastering Diverse Domains through World Models
- PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
- Causal World Modeling for Robot Control
- FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models
- WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos
- EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations
- LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion
- DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
- R3M: A Universal Visual Representation for Robot Manipulation
- DINOv2: Learning Robust Visual Features without Supervision
- EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World
- Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
- HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos
- Flow as the Cross-Domain Manipulation Interface
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving