UNITAS: A 3D-Native World Action Model for Embodied Manipulation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "UNITAS: A 3D-Native World Action Model for Embodied Manipulation".
Dev: The gist The first 3D-native world action model that unifies observations, actions, and scene dynamics in a shared metric 3D frame within each interaction,
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: We're talking about "UNITAS: A three dee-Native World Action Model for Embodied Manipulation" and how it fundamentally changes how we model robot interaction in space <ref:2610.12099#pg1,UNITAS: A 3D-Native World Action Model for Embodied Manipulation>.
Dev: This paper is focused on creating this first model that unifies observations, actions, and scene dynamics all within a shared metric three dee frame for every interaction <ref:2610.12099#pg1,model that unifies observations, actions, and scene dynamics>.
Taro: What’s really interesting is that they tackle the problem of representation because existing models usually rely on images or visual latents which are view-dependent projections.
Rosa: So UNITAS introduces a completely different way to handle this by using three dee point trajectories as the action flow and scene flow, which directly encode physical distances <ref:2610.12099#pg1>.
Dev: They achieve this by grounding observations with world-aligned three dee positional embeddings, which allows them to work even when you don't have depth input <ref:2610.12099#pg1>.
Taro: That grounding mechanism is paired with a physical-time trajectory tokenizer that turns the motion history into tokens anchored at their current three dee position <ref:2610.12099#pg1>.
Rosa: This means the model has a common representation across different robot embodiments, which is something that’s been really hard to get before.
Dev: The paper highlights how they achieve this by using a common interface for both policy mode and simulator mode operation, depending on what you need right then.
Taro: It seems like they are explicitly making the interaction geometry an explicit prediction target, which is a pretty novel way to connect world modeling with action learning.
Rosa: And the performance numbers are strong, hitting ninety-nine point eight percent success on LIBERO in policy mode and showing solid gains in simulator metrics like reducing Moving ADE by up to forty-nine percent <ref:2610.12099#pg2>.
Dev: The caveat they mention is that ablation studies show that adding scene-flow supervision to the reference model really helps boost robustness, which tells us more about what’s needed for a reliable system.
The paper's summary: Rosa: To summarize, UNITAS is this three dee-native world action model that uses point trajectories to represent both human hands and robot grippers as its action flow <ref:2610.12099#pg1,3D-native world action model>.
Dev: This action flow then conditions the scene flow, which describes the environmental response based on those motions, all within a shared metric frame.
Taro: The architecture is built around these tokens: observation tokens grounded in world coordinates, action-flow tokens across embodiments, and scene-flow tokens anchored at their current three dee position <ref:2610.12099#pg1>.
Rosa: The training setup uses two diffusion transformers—one for the dynamics of the action and scene flows, and another one for the end-effector action chunk u.
Dev: The objective function L = λALA + λsLscene + λctrlLctrl + λPELPE uses mean squared error losses to supervise all those components.
Taro: The core contribution is establishing this unified representation, showing that we can align observations, actions, and scene dynamics in one consistent three dee space <ref:2610.12099#pg1,observations, actions, and scene dynamics in>.
Rosa: It’s about making the interaction geometry a central part of the model's prediction targets so it learns how to move and how the world reacts together.
Dev: This structure allows for direct action execution without needing to generate action or scene flows separately, which is efficient during inference.
Taro: So, fundamentally, they are showing that you can achieve high manipulation results by explicitly modeling the physical relationship between the robot and its environment in three dee space <ref:2610.12099#pg1>.
The paper's improvements: Rosa: One major improvement they highlight is using world-aligned three dee positional embeddings to ground visual observations, which lets them work with or without depth input <ref:2610.12099#pg1>.
Dev: That feature is important because it makes the perception more robust, which is a big win in real-world scenarios where depth sensors might not always be perfect.
Taro: They also introduced the physical-time trajectory tokenizer, encoding each point trajectory as one token anchored at its current three dee position <ref:2610.12099#pg1>.
Rosa: That tokenizer is clever because it lets the model represent motion across different frame rates and durations uniformly by anchoring it spatially.
Dev: The paper shows that world-aligned visual grounding is crucial for policy performance and robustness, because when you remove that encoding, success drops from seventy-nine point one one percent to sixty-six point seven three percent under camera perturbations <ref:2610.12099#pg2>.
Taro: And they also showed that adding scene-flow supervision to the reference model improves overall success from eighty-five point seven four percent up to eighty-six point four seven percent, which points toward a direction for future work in supervised learning methods <ref:2610.12099#pg2>.
Rosa: So these improvements show that having consistent physical grounding and supervisory signals across action flow and scene flow really improves the policy's ability to handle changes in the environment.
Dev: The paper notes that their system can perform zero-shot scene-flow predictions on real observations conditioned on ground truth gripper trajectories, suggesting those learned interaction dynamics actually transfer from simulation to real scenes.
Conclusion: Rosa: So wrapping up UNITAS, the main implication is that we have a model that consistently aligns observations, actions, and scene dynamics in a shared metric three dee frame <ref:2610.12099#pg1,observations, actions, and scene dynamics in a shared metric 3D frame>.
Dev: This unified representation across robot embodiments and human hands is what makes it so powerful for manipulation tasks.
Taro: It moves the field by making interaction geometry an explicit prediction target, which connects world modeling directly with action learning in a way that's hard to do otherwise.
Rosa: It sets a strong foundation for embodied learning by representing robot behavior and environmental change as complementary parts of the same process.
Dev: The performance on real hardware, hitting ninety percent success on Organize Books, shows this formulation has real value for deployment outside of the lab <ref:2610.12099#pg2>.
Taro: I just want to say that while they show strong simulation results, extending this foundation to reliable prediction over longer interactions and broader real-world deployment is where the next big step lies.
Rosa: Agreed, it’s a solid model to build on as we push for more general purpose robot intelligence.
Dev: Good talk today on UNITAS. We'll be back after the break with another paper from arXiv.
Ruixiang Wang, Yongyi Su, Wenlve Zhou, Bo Yue, Hengyan Liu, Dekun Lu, Yuxin Tian, Yihan Fang, Zerui Wu, Xing Hu, Jietao Chen
The Chinese University of Hong Kong, Shenzhen
cs.RO
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/DexForce/UNITAS
The gist: The gist The first 3D-native world action model that unifies observations, actions, and scene dynamics in a shared metric 3D frame within each interaction, using a common representation across robot
Key concepts
- Action Flow
- This concept represents the movement of human hands or robot grippers as continuous 3D point trajectories. These trajectories serve as a common way to describe physical interaction across different robot bodies and human hands, providing a unified representation for action.
- Scene Flow
- Scene flow describes how points in the environment move or change their position based on the action flow. It is decoded from tokens anchored at specific 3D positions, allowing the model to predict environmental displacements conditioned directly on the planned motion.
- World-Aligned Positional Embeddings
- These embeddings ground visual observations with or without depth information directly into world coordinates. This ensures that all visual inputs are consistently mapped to a shared metric 3D frame, which is crucial for aligning observations, actions, and scene dynamics.
- Physical-Time Trajectory Tokenizer
- This component encodes each individual point trajectory (like a hand movement) as a single token. This token is spatially anchored at the current 3D position of that point in the world, creating a structured interface that connects physical motion directly to model tokens.
Terminology
Summary
The gist The first 3D-native world action model that unifies observations, actions, and scene dynamics in a shared metric 3D frame within each interaction, using a common representation across robot embodiments and human hands.
How it works
UNITAS is introduced as the first 3D-native world action model that aligns observations, actions, and scene dynamics in a shared metric 3D frame within each prediction window, using a common representation across robot embodiments and human hands It achieves this by representing action flow as human hands and robot grippers as 3D point trajectories, while scene flow describes scene-point displacements conditioned on these trajectories World-aligned 3D positional embeddings ground visual tokens with or without depth input Furthermore, a physical-time trajectory tokenizer encodes each point trajectory as one token anchored at its current 3D position This interface supports both direct action execution and action-conditioned scene prediction.
Key Components and Representations
The model factorizes the joint distribution of action flow, executable actions, and scene flow as pθ(A, u, S C, q0) = pθ(A C) pθ(u C, q0) pθ(S A, C). Action flow represents human hands and robot grippers as 3D point trajectories Scene flow describes scene-point displacements decoded from trajectory tokens conditioned on action flow. The backbone of UNITAS contains observation tokens grounded in world coordinates, while action-flow tokens across embodiments and scene-flow tokens are anchored at each point’s current 3D position.
Training and Objectives
The model is trained using a Mixture-of-Transformers with two diffusion transformers: a point dynamics expert for action and scene flows, and an action expert for the end-effector action chunk u. The training objective is L = λALA + λsLscene + λctrlLctrl + λPELPE. The first three terms use mean squared error (MSE) losses to supervise action flow, scene flow, and the action chunk. Scene prediction conditions on clean action-flow latents from demonstrations during training.
Performance and Evaluation
UNITAS achieves leading manipulation results among compared methods with 1.7B parameters. In simulator mode, it leads compared methods across all four RoboTwin scene-prediction metrics, reducing Moving ADE by 35%, 21%, and 49% relative to PointWorld. In policy mode, it achieves 99.8% success on LIBERO and outperforms larger video-based WAMs on RoboTwin. Real-robot success reaches 85% across real-world tasks.
Contributions
The main contributions are:
• To our knowledge, UNITAS is the first 3D-native world action model that aligns observations, actions, and scene dynamics in a shared metric frame, using a common representation across robot embodiments and human hands.
• Action flow represents human hands and robot grippers through point trajectories, while scene flow represents the environmental response conditioned on that motion.
• World-aligned positional embeddings ground visual observations with or without depth, and a physical-time trajectory tokenizer encodes each point trajectory as one token spatially anchored at its current 3D position.
• With 1.7B parameters, UNITAS achieves leading action-conditioned scene prediction and strong manipulation performance in simulation and on real hardware.
Inference Efficiency
UNITAS directly produces executable actions without generating action or scene flows in 55.77 ms per chunk, providing a 2.20× speedup over jointly generating actions and action flow. In world-model mode, the number of scene queries can be adjusted at inference to match the desired spatial sampling density and compute budget. Reducing scene queries from 3,000 to 1,024 lowers joint prediction latency from 561.60 ms to 305.49 ms.
Real-World Transfer
The model demonstrates zero-shot scene-flow predictions on real-world observations conditioned on ground-truth gripper trajectories, suggesting that the learned interaction dynamics can transfer from simulation to real scenes. This qualitative result suggests that the learned interaction dynamics can transfer from simulation to real scenes.
Ablation Study Insights
Ablations show that world-aligned visual grounding, action- and scene-flow supervision, and motion history improve policy performance and robustness. Removing 3D positional encoding reduces success to 81.35%, with the largest drop under camera perturbations, from 79.11% to 66.73%. Adding scene-flow supervision to the reference model improves overall success from 85.74% to 86.47%, suggesting that future scene preTable. The results support positive cross-dataset transfer under our unified modeling formulation.
Conclusion
UNITAS advances a physical-space formulation of world action modeling: what the robot observes, how it moves, and how the scene responds are expressed within a common metric reference. Making interaction geometry an explicit prediction target connects world modeling with action learning and provides a shared representation across differing embodiments and observation views. The strong policy performance observed in simulation supports the value of this formulation for manipulation. More broadly, this work positions physical interaction as a unifying basis for embodied learning, with robot behavior and environmental change represented as complementary parts of the same process. Extending this foundation to reliable prediction over longer interactions and broader realworld deployment is a promising direction for general-purpose robot intelligence.
--- Page 1 ---
UNITAS is a 3D-native world action model that unifies observations, actions, and scene dynamics in a shared metric 3D frame within each interaction, using a common representation across robot embodiments and human hands It achieves this by representing action flow as human hands and robot grippers as 3D point trajectories, while scene flow describes scene-point displacements conditioned on these trajectories.
How it works
World-aligned 3D position embeddings ground visual tokens with or without depth input A physical-time trajectory tokenizer encodes each point trajectory as one token anchored at its current 3D position. This interface supports both direct action execution and action-conditioned scene prediction.
Key Components and Representations
Action flow represents human hands and robot grippers as 3D point trajectories Scene flow describes scene-point displacements decoded from trajectory tokens conditioned on action flow. The backbone of UNITAS contains observation tokens grounded in world coordinates, while action-flow tokens across embodiments and scene-flow tokens are anchored at each point’s current 3D position.
Training and Objectives
The model is trained using a Mixture-of-Transformers with two diffusion transformers: a point dynamics expert for action and scene flows, and an action expert for the end-effector action chunk u. The training objective is L = λALA + λsLscene + λctrlLctrl + λPELPE. The first three terms use mean squared error (MSE) losses to supervise action flow, scene flow, and the action chunk<ref:2610.
Improvements for AI systems
-
World-aligned 3D positional embeddings ground visual tokens
with or without depth input,
enabling robust perception even when depth is unavailable, as shown in Equation 6 and Appendix B. -
The physical-time trajectory tokenizer encodes each point trajectory
as one token anchored at its current 3D position,
allowing the model to represent motion across different frame rates and durations uniformly. -
UNITAS supports both
direct action execution and action-conditioned scene prediction
by factorizing the joint distribution aspθ(A, u, S C, q0) = pθ(A C) pθ(u C, q0) pθ(S A, C),
enabling both policy mode and simulator mode operation. -
The model achieves superior manipulation success on real-world tasks by learning a
shared metric 3D frame within each prediction window
and using acommon representation across robot embodiments and human hands.
-
The system can perform zero-shot scene-flow predictions in real environments, as demonstrated in Figure 11, by leveraging learned interaction dynamics that
can transfer from simulation to real scenes
despite lacking explicit scene-trajectory ground truth.
Sources
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
- Motus: A Unified Latent Action World Model
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- WorldVLA: Towards Autoregressive Action World Model
- LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
- ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots
- Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow
- RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
- PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- OpenVLA: An Open-Source Vision-Language-Action Model
- Causal World Modeling for Robot Control
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving