From World Models to World Action Models: A Concise Tutorial for Robotics
cs.RO, cs.AI, cs.SY, eess.SY
Submitted: 2026-07-01
Updated: 2026-09-08
Comments: Github page: https://github.com/clearlab-sustech/WorldModelSurvey
Code: https://github.com/clearlab-sustech/WorldModelSurvey
License: http://creativecommons.org/licenses/by/4.0/
The gist: Rather than providing an exhaustive survey, this paper presents a concise tutorial on world models and world action models for robotics.
Terminology
Abstract
Rather than providing an exhaustive survey, this paper presents a concise tutorial on world models and world action models for robotics. After reading the tutorial, readers should have a clear understanding of what constitutes a "world", how world models and world action models are defined, and what roles they play within robotic AI systems. The tutorial also develops a unified perspective for comparing representative approaches, such as World Labs' spatial intelligence models, Yann LeCun's JEPA framework, and NVIDIA's Cosmos platform, and clarifies how these models differ in their representations, predictive capabilities, and interaction mechanisms.
Sources
- Learning Universal Policies via Text-Guided Video Generation
- Vidar: Embodied Video Diffusion Model for Generalist Manipulation
- AdaWorld: Learning Adaptable World Models with Latent Actions
- DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
- OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation
- Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models
- Mastering Diverse Domains through World Models
- GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
- ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation
- VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning
- PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
- ExoPredicator: Learning Abstract Models of Dynamic Worlds for Robot Planning
- LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation
- Grounding Video Models to Actions through Goal Conditioned Exploration
- Object-Centric World Model for Language-Guided Manipulation
- DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
- PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving