CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
cs.RO, cs.AI, cs.CV
Submitted: 2026-08-27
Updated: 2026-09-11
Code: https://github.com/omni-CLAP/clap
Project page: https://omni-clap.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich
Terminology
Abstract
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents. CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actor. However, cross-embodiment learning is non-trivial because action representations vary sharply across robot platforms and are typically absent in human videos. CLAP addresses this fundamental challenge through the following core contributions. First, CLAP reconciles disparate action spaces using end-effector poses, language instructions, and latent actions. Second, to resolve their individual limitations, CLAP introduces a curriculum-based cross-embodiment learning recipe that first learns foundational physical priors across unlabeled video data using latent actions and subsequently grounds them in end-effector action spaces for zero-shot deployment to real-world tasks. Crucially, CLAP approaches or surpasses state-of-the-art single-embodiment video models in challenging environments like DROID. These performance advantages compound via few-shot adaptation to establish a novel paradigm for training single-embodiment video world models. Ultimately, CLAP delivers the most comprehensive suite of action-conditioned video world models to date - spanning diverse action-conditioning spaces (end-effector, language, and latent) and robot morphologies (including cross-embodiment, DROID, Bridge, bimanual YAM robots, and G1 humanoids). We open-source all code and models. Project Website at https://omni-clap.github.io.
Sources
- LLaMA: Open and Efficient Foundation Language Models
- PaLM: Scaling Language Modeling with Pathways
- Scaling Laws for Neural Language Models
- Language Models are Few-Shot Learners
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- WorldGym: World Model as An Environment for Policy Evaluation
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- MolmoAct2: Action Reasoning Models for Real-world Deployment
- Cosmos World Foundation Model Platform for Physical AI
- Wan: Open and Advanced Large-Scale Video Generative Models
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models
- Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation
- Video Language Planning
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
- Robotic Video World Models: A Survey of Applications, Research Challenges, Future Directions
- IRASim: A Fine-Grained World Model for Robot Manipulation
- Scalable Policy Evaluation with Video World Models
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving