Motus2: A Self-Evolving General World Model for Dexterous Manipulation
cs.RO, cs.AI, cs.CV, cs.LG
Submitted: 2026-08-31
Updated: 2026-09-10
Project page: https://motus-robotics.github.io/motus2
License: http://creativecommons.org/licenses/by/4.0/
The gist: General embodied agents should perceive, predict, act, evaluate, and improve within a unified system.
Terminology
Abstract
General embodied agents should perceive, predict, act, evaluate, and improve within a unified system. World models have shown great promise in building such agents, yet existing models typically append an action output head to a world simulator, without coupling them into a closed decision-and-learning loop for policy improvement. We present Motus2, a self-evolving general world model for dexterous manipulation. Motus2 advances world modeling through model scaling and data scaling. For model scaling, a single model with shared weights exposes three control interfaces: a policy (world-action model), a simulator (action-conditioned world model), and an evaluator (value model). The policy proposes candidate action chunks, the simulator predicts their visual consequences, and the evaluator assesses the predicted outcomes. Their coupling forms a closed decision-and-learning loop for policy improvement. This formulation uses curated expert demonstrations for action learning, while failed and suboptimal interactions provide valuable evidence for dynamics modeling and value learning. For data scaling, Motus2 progresses from large-scale monocular egocentric data to synchronized stereo egocentric data, followed by robot-domain adaptation with robot trajectories and supplementary human-robot alignment data. Motus2 further studies global-autoregressive and hybrid-memory extensions of its sliding-window context, adds tactile feedback for contact-aware control, and is instantiated on a fully biomimetic platform with stereo vision, dual arms, dual dexterous hands, and tactile sensing. Together, egocentric data scaling and closed-loop general world model scaling provide a general path toward self-evolving dexterous manipulation.
Sources
- Motubrain: An Advanced World Action Model for Robot Control
- Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
- T-Rex: Tactile-Reactive Dexterous Manipulation
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- MemoryWAM: Efficient World Action Modeling with Persistent Memory
- RELIC: Interactive Video World Model with Long-Horizon Memory
- Mastering Diverse Domains through World Models
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
- GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
- METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
- EgoSteer: An Open-Source Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos
- EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data
- Reinforcing Action Policies by Prophesying
- NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards
- Wan: Open and Advanced Large-Scale Video Generative Models
- GigaWorld-Policy: An Efficient Action-Centered World--Action Model
- AcceRL: A Distributed Asynchronous Reinforcement Learning and World Model Framework for Vision-Language-Action Models
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving