Cross-Embodiment Robot Foundation World Models with Latent Actions
summary
The gist
The gist: The Latent Action Conditioned Robot World Model (LAC-WM) improves downstream performance over an Explicit Action-Conditioned World Model (EAC-WM) by up to 46.7% on dexterous manipulation
In short
The Latent Action-Conditioned Robot World Model (LAC-WM) was introduced to improve robot world models across different robot bodies. It uses a unified latent action space shared by diverse embodiments, which outperforms explicit action conditioning. This unified space allows performance to scale positively with the number of pretraining robots, unlike explicit methods which degrade.
Key concepts
- Latent Action-Conditioned Robot World Model (LAC-WM)
- This model operates using a single, shared latent action space that all robot bodies can use. It abstracts away physical differences so the model learns general control skills applicable to any robot, making it highly adaptable when faced with new robot designs.
- Explicit Action-Conditioned World Model (EAC-WM)
- This method conditions the world model directly on explicit motion labels—specific commands for movement. The paper shows this creates separate action representations for each robot body, which hinders the model's ability to generalize when switching to a new robot.
- Unified Latent Action Space
- This is a shared mathematical space where actions from different robots are mapped. By forcing all actions into one space, the model can effectively integrate data from many different embodiments during pretraining, leading to better overall performance and scalability.
- Scaling Law with Number of Embodiments
- This analysis shows how well the world model performs as more robot bodies are used for training. LAC-WM's performance improves as more robots are included, while EAC-WM's success rate decreases, proving the unified space is crucial for efficient cross-embodiment learning.
Terminology used across episodes
This episode discusses
- Cross-Embodiment Robot Foundation World Models with Latent Actions · Paper Radio
- Cosmos World Foundation Model Platform for Physical AI
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI
- AdaWorld: Learning Adaptable World Models with Latent Actions
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
- GAIA-1: A Generative World Model for Autonomous Driving
- pi 0.5: a Vision-Language-Action Model with Open-World Generalization
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models
- Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
- CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- Long-Context State-Space Video World Models
- GameFactory: Creating New Games with Generative Interactive Videos
- Matrix-Game: Interactive World Foundation Model
- Universal Actions for Enhanced Embodied Foundation Models
The paper
Cross-Embodiment Robot Foundation World Models with Latent Actions · Read on arXiv
Huang Huang, Sriram Yenamandra, Arjun Majumdar, Elie Aljalbout, Tushar Nagarajan, Tsung-Yen Yang, Akshara Rai, Michael Rabbat
Stanford University
The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments. This unified action space improves the world model's performance when adapted to previously unseen robot embodiments. We compare LAC-WM with an Explicit Action-Conditioned World Model (EAC-WM), which conditions on explicit motion labels. Our results show that explicit action conditioning leads to disjoint action representations across embodiments, limiting downstream performance when adapting to new robots. We evaluate both models on dexterous manipulation tasks and a modified LIBERO benchmark. LAC-WM improves downstream performance over EAC-WM by up to 46.7% on dexterous manipulation and 11.7% on LIBERO. Crucially, the unified latent action space allows LAC-WM's downstream performance to scale positively with the number of embodiments used during pretraining. In contrast, the disjoint action space in EAC-WM leads to decreased performance as the number of pretraining embodiments increases. These results highlight the importance of a unified action space for efficient cross-embodiment learning, addressing a key challenge in robotics. Project website: https://lacwm.github.io/
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Cross-Embodiment Robot Foundation World Models with Latent Actions".
Dev: The gist:
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So what exactly does this paper suggest is the main contribution here? Dev Well, it's about moving away from models that rely on explicit action labels and towards a unified latent action space. Rosa Right, so instead of teaching the model what every single robot does separately, they are creating one shared language for actions across different robots.
Taro: It’s like they’re trying to find that common ground in how robots move, even if their physical setups look completely different. Dev That unified space is key because it lets the world model adapt better when you try to use a robot you haven't seen before during finetuning.
Rosa: It says this unified action space improves performance when the model has to deal with previously unseen robot embodiments. Taro That’s what they’re showing, that the unified space helps them generalize across different physical setups without losing information about the task itself.
Dev: They compare this against an Explicit Action-Conditioned World Model, or EAC-WM, which relies on explicit motion labels for conditioning. Rosa And the result there is pretty telling: explicit action conditioning leads to disjoint action representations across embodiments.
Taro: So if you train a model that only knows how Robot A moves, it struggles when you show it Robot B because their actions don't map onto the same space. Dev That’s the core issue they are pointing out with EAC-WM.
The paper's summary: Rosa: So they summarize this by saying that LAC-WM improves downstream performance significantly over EAC-WM, showing gains of up to forty-six point seven percent on dexterous manipulation tasks and eleven point seven percent on a benchmark called LIBERO <ref:2610.10846#pg1,up to 46.7% on dexterous manipulation>. Dev Those are pretty big numbers for real-world robot skills, especially the manipulation part.
Taro: The summary emphasizes that this improvement comes directly from the unified latent action space, which allows performance to scale positively with the number of embodiments used during pretraining. Rosa So, if you train it on ten different robots, it gets better at handling a new robot than if you only trained it on one specific robot.
Dev: The paper explains that LAC-WM has four parts: an inverse dynamics model, which encodes observations into those latent actions, a forward dynamics model to predict the next frame based on them, a motion decoder to translate those actions back into labels, and an action projector. Rosa It’s a whole pipeline working together in the pretraining phase.
Taro: And they also point out they introduced an auxiliary motion decoding loss which helps align those latent actions with explicit control actions, which is designed to make sure the latent space captures actual actionable content. Dev That’s their way of encouraging the model to learn something physically meaningful in that shared space.
The paper's improvements: Rosa: What are the specific technical improvements they highlight? Taro They show how this unified latent action space allows for semantic consistency in video generation, which is really cool because it means the same latent action prompt can move both a robot’s end-effector and a human’s hand in the same direction. Dev It suggests that their learned latent actions are capturing movement concepts that are shared across different physical forms.
Rosa: They also show this space is much better at capturing motion than the disjoint spaces used in EAC-WM, which makes video generation look more coherent when you switch embodiments. Taro And for planning, they say world-model–based action selection significantly outperforms VLA-only baselines, achieving a twenty-nine point zero percent higher success rate on the S.R metrics compared to the strongest VLA baseline they tested.
Dev: They also look at how it scales when you increase the number of pretraining embodiments used; LAC-WM’s downstream planning performance improves as that number increases, whereas EAC-WM shows the opposite trend with decreased success rates. Rosa That scaling behavior is a really important part of their argument for why this unified approach is better for foundation models.
Conclusion: Rosa: To wrap up, the paper makes a strong case that conditioning on this unified latent action space enhances cross-embodiment learning and adaptation to new robots compared to using a disjoint latent space. Dev Basically, the main point is that sharing one action language across all embodiments helps build more robust robot foundation models for planning.
Taro: They also mention that they found that both the motion decoding loss and crossaugmentation inputs are important when looking at downstream robot planning success rates. Rosa So what this means for us who just listen to the show is that if you’re building a system to control robots, thinking about a shared action space instead of separate ones might be the way forward.
Dev: The paper "Cross-Embodiment Robot Foundation World Models with Latent Actions" shows that by focusing on learning a unified latent action space, we can build world models that are much more adaptable and perform better when we move from lab training to real-world deployment across different hardware.
Rosa: Yeah, so they’ve laid out a clear path for how to make these large robot models actually useful in varied environments. We'll see what the next set of papers brings.
More episodes
- 2610.10601-Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection
- 2610.10637-TacHair: Tactile Contact-Distribution Guided Online Correction for Robotic Hair Stroking and Perception
- 2610.10646-Masked Generative Motion Planning with Geometry-Guided Token Search
- 2610.10812-Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation
- 2610.10801-Same Action, Different Outcome: Variability in Dynamic Cloth Manipulation
- 2610.10810-Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams
- 2610.10748-TAPNAV: Humanoid Navigation through Tactile Active Perception
- 2610.10855-OmniHOI: Dexterous Hand-Object Interaction from Monocular Human Video
- 2610.11003-ActiveReg: Information-Driven Active Regional Probing for Partial-to-Full Bone Registration
- 2610.10905-Informationally Decoupled Trajectory Design for Sim-to-Real System Identification