World Models for Embodied Intelligence: From Plausible to Controllable to Actionable
cs.RO, cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: Project Page: https://3dagentworld.github.io/EmbodiedWM/
Project page: https://3dagentworld.github.io/EmbodiedWM
License: http://creativecommons.org/licenses/by/4.0/
The gist: World models connect perception and decision-making in embodied intelligence by maintaining hidden state, anticipating consequences, comparing interventions, and adapting when execution departs from
Terminology
Abstract
World models connect perception and decision-making in embodied intelligence by maintaining hidden state, anticipating consequences, comparing interventions, and adapting when execution departs from expectations. Although progress is often measured by visual fidelity, their value lies in improving behavior. Before reaching for a cup, a person anticipates its weight and resistance to grasping, shaping the hand before contact. Such anticipation is coarse and rarely pictorial, yet it guides action. This raises a central question: which predictive capabilities improve behavior? Existing surveys, organized by architecture, output modality, or application domain, leave this question implicit. We introduce three progressively stronger capability levels: Plausible models preserve task-relevant temporal, geometric, or physical structure; Controllable models additionally predict how interventions alter that structure; and Actionable models translate predictions into measurable gains in planning, action, learning, evaluation, verification, recovery, or data selection. We complement this hierarchy with a 3 x 4 matrix crossing geometry, physics, and action grounding with improvement loops centered on data, rewards, policies, and the model itself. Using this framework, we survey manipulation, navigation, locomotion, autonomous driving, and general embodied learning, tracing technical progressions, clarifying capability requirements, and examining datasets, benchmarks, and evaluation protocols. We identify challenges in long-horizon consistency, uncertainty calibration, causal intervention testing, latency, verification and recovery, and cross-embodiment transfer. This perspective shifts evaluation from visual plausibility toward whether predictions capture task-relevant state, reflect intervention effects, and improve the closed-loop behavior of embodied agents.
Sources
- Physically Embodied Gaussian Splatting: A Realtime Correctable World Model for Robotics
- Cosmos World Foundation Model Platform for Physical AI
- Feedback World Model Enables Precise Guidance of Diffusion Policy
- ACT-Bench: Towards Action Controllable World Models for Autonomous Driving
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- FitVid: Overfitting in Pixel-Level Video Prediction
- Walk through Paintings: Egocentric World Models from Internet Priors
- AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing
- WorldVLA: Towards Autoregressive Action World Model
- GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
- GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
- Large Video Planner Enables Generalizable Robot Control
- TransDreamer: Reinforcement Learning with Transformer World Models
- Goal-VLA: Image-Generative VLMs as Object-Centric World Models Empowering Zero-shot Robot Manipulation
- LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
- Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control
- ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment
- RegGS: Unposed Sparse Views Gaussian Splatting with 3DGS Registration
- Graph-Guided Scene Reconstruction from Images with 3D Gaussian Splatting
- Unposed 3DGS Reconstruction with Probabilistic Procrustes Mapping
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving