Humanoid World Action Model With Joint State--Action Generation

arXiv:2610.12026 · cs.RO, cs.AI · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Humanoid World Action Model With Joint State--Action Generation".

Dev: The gist The HWAM introduces a Humanoid World Action Model with joint state–action generation,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So we're looking at this paper today, "Humanoid World Action Model With Joint State--Action Generation," by Yan Yang and team. It's about how humanoid robots can bridge that gap between what the brain wants to do and what the robot actually does when it moves.

Dev: Yeah, it tackles that action-execution gap head-on. Basically, they introduce a model where the robot’s state after an action is explicitly used as a target for generating the next set of actions.

Taro: So if I'm driving and I tell the car to turn left, but it actually drifts right because of some weird physics happening under the hood, this paper tries to figure out how to make sure we model that drift.

Rosa: Exactly. The core idea is defining a joint state-action target, Y t = Concat(S t, A t), where S t is the state and A t is the action at time h, which they show on page three of this paper. This means both channels are denoised together by a shared generator.

Dev: That joint denoising thing is key because it gives the action generation part access to the evolving state estimate throughout the entire denoising process, which helps connect the high-level policy to what's physically happening right now.

Taro: It sounds like they are trying to make sure that whatever reference action you give is grounded in what's actually happening physically on the robot at every step.

Rosa: Right. And they don't just stop there; they connect this joint target into three different training paths for better learning, which is a big part of this paper.

Dev: That’s where they introduce the Policy path, which only denoises state-action trajectories conditioned on current observations, matching the deployment conditions we see in real life.

Taro: So that means the policy learns to generate actions that are relevant to how the robot is actually set up when it's running.

Rosa: Then they have Forward Dynamics Modeling, or FDM, which predicts future visual observations by conditioning them on both the actions and those post-execution states they just calculated.

Dev: That explicitly accounts for the realized body motion when predicting scene changes, so it’s not just guessing where the robot will go based on the command.

Taro: And then there’s Inverse Dynamics Modeling, or IDM, which goes in reverse by reconstructing the joint trajectory from visual transitions to supervise actions and states they produce.

Rosa: It sounds like a bidirectional grounding mechanism where you can learn how to get there and learn how to get back from the visual results.

Dev: They connect all those paths using path-dependent masked cross-attention between a video expert and their state-action expert, using conditional flow matching objectives for both modalities.

Taro: So they’re mixing visual prediction with physical dynamics modeling through these different training routes to get a robust understanding of the system.

Rosa: And when we look at the results, they show this model achieving a seventy point six percent success rate on Candy Picking compared to forty-three point three percent for Fast-WAM on the LimX OLI humanoid robot.

Dev: That's a significant jump in success under those real-robot tasks, which is what we really want to see when moving from simulation to actual hardware.

Taro: I also saw some interesting ablation data where adding this joint target improved success from fifty-five point zero percent to seventy-three point three percent when using FDM, IDM, and Policy training in mobile manipulation tasks.

Rosa: That tells us that the benefit of this joint state-action generation depends on how you train the system; it helps a lot when you use forward dynamics and inverse dynamics together with the policy path.

Dev: But they also found that under VGM and Policy training, adding the joint target actually decreased success from sixty point zero percent down to forty-five point five percent, which shows that we have to be careful about which training recipe we pick for a given setup.

Taro: So it’s not a magic fix; it’s about matching the right components—the state-action target, the FDM, and the IDM—to the specific learning environment.

Rosa: That brings us to wrapping up this paper on "Humanoid World Action Model With Joint State--Action Generation." It successfully introduces joint state-action generation to address that action-execution gap in hierarchical humanoid systems.

Dev: It achieves high success rates on the evaluated OLI tasks, and the ablation studies give us some important guidance on how to configure the training paths for better results.

Taro: For me, what’s really interesting is that this method doesn't try to explicitly identify low-level dynamics; instead, these predicted states are treated as jointly generated variables connecting high-level actions to their physical and visual consequences.

Rosa: So in short, it’s a way to connect the high-level policy directly to the physical reality of motion by making the post-execution state an explicit prediction target.

Dev: It's about using those forward and inverse dynamics models alongside a policy path that learns from observations to create a more grounded action reference.

Taro: I think this means future work will focus on how this structure holds up when the environment is much messier or when the humanoid has different physical characteristics than what they tested on.

Rosa: Right, so we’ll keep an eye on how this HWAM performs in more complex, unconstrained real-world scenarios.

Dev: We've got a lot to unpack there about those training recipe dependencies next time.

The paper's summary: Rosa: So, to recap, this new model connects the high-level action you’re planning directly to what actually happens on the robot's body by making both its current state and its intended action a single target for training.

Dev: Yeah, it’s about solving that disconnect where the command you give doesn't quite match the motion you see because of how physics works. It explicitly forces them to learn how those two things—the state and the action—are related together in a shared way.

Taro: What this means for me is that when a robot messes up, it doesn't just get penalized for the wrong action; it gets penalized for having a state that didn't match what the action promised. It forces the model to build a more physically grounded understanding of its movements.

Rosa: Right. The paper sets up three different ways to train this system, and they use all three together—a policy path that learns from what it sees right now, forward dynamics modeling which predicts future visuals based on that state, and inverse dynamics modeling which tries to figure out the state from the visual changes.

Dev: That’s where the real meat is. It’s not just one training method; it’s a whole system of interconnected learning paths. The policy path ensures deployment conditions are matched, while FDM and IDM create a bidirectional link between what it wants to do and what it actually sees visually.

Taro: And the results show that this combined approach is much more robust. They found that when you use the forward dynamics and inverse dynamics together with the policy training, the success rate on those real-world tasks like Candy Picking jumps quite a bit.

Rosa: It does, specifically hitting seventy point six percent success on that humanoid task compared to lower numbers from other methods they tested. The paper also shows how this joint target helps even when you change your training recipe—it performs better in one setup but less well in another, which tells us a lot about how the learning process itself influences the final result.

Dev: That dependence on the training recipe is important for me, because it means we can’t just throw this model at every problem and expect it to work perfectly everywhere. It depends on whether you're using forward dynamics or inverse dynamics as your main supervisor.

Taro: So, if I only listen to the numbers, it suggests that getting that state-action connection right is crucial for complex manipulation, but you have to tune the training setup carefully depending on the robot and the task.

Rosa: Exactly. And what this implies for us is that we need these kinds of models when we move from lab simulations to real-world environments where things are messy and unpredictable.

Dev: Because if you can ground a high-level plan in physical reality like this, it means the AI isn't just making guesses; it’s building something that understands the actual physics of moving.

The paper's improvements: Tom: So, to wrap up the core idea, they are proposing a joint state-action target that lets both parts of the model learn from each other during training.

Rosa: Exactly. It’s about forcing the state and action channels to be denoised together because they are physically linked in execution.

Dev: And this specific target structure, Y t = Concat(S t, A t), means the action generator gets access to the evolving state estimate throughout the whole process.

Taro: It’s a big step because it means when you're doing a task with an AI robot, you’re not just telling it what to do; you’re giving it a target that includes *how* it ended up there physically.

Rosa: Right. They then connect this joint target into three different training paths—policy, forward dynamics modeling, and inverse dynamics modeling—to make sure the robot learns both how to plan and how to execute realistically.

Dev: The main improvement is that the system can learn from execution failures in a way that respects the physical constraints of the robot's body motion. It’s not just about matching an output; it’s about matching the *process* of getting there.

Taro: I think this structure helps when things go wrong in complex environments, because it provides supervision both forward and backward through those dynamics models.

Rosa: And they showed that this joint target helps when you train with forward and inverse dynamics alongside the policy path, improving success on tasks like Candy Picking. seventy point six percent is a solid number for real-robot work.

Dev: But they also found that it can actually hurt performance if you use a different training setup, showing that the choice of training recipe matters a lot. You have to be smart about which dynamics model you pair with the policy path for best results.

Taro: That means we can’t just assume one way works for every robot; we have to adapt our whole learning strategy based on what we’re trying to achieve.

Rosa: So, the implication is that this method gives us a stronger framework for humanoid systems by explicitly linking the high-level plan to the low-level physical reality, even if it means tuning the training setup for every new robot.

Conclusion: Rosa: So to wrap up, this paper introduces the Humanoid World Action Model With Joint State--Action Generation, which uses that joint target to explicitly link what an AI robot wants to do with what it actually does physically.

Dev: Yeah, essentially it tackles that action-execution gap by making the post-execution state a shared prediction target for both channels. It’s a lot of work on the loop rate and latency side because you’re tying them together so tightly.

Taro: What this means for autonomy is that we can start modeling not just the command, but also the resulting physical consequence right away, which is crucial when the world doesn't behave exactly as expected.

Rosa: Right. They showed it works on real-robot tasks like Candy Picking with a seventy point six percent success rate, which is pretty solid for humanoid manipulation outside of just simulation.

Dev: That number is impressive considering the complexity of whole-body dynamics involved in those kinds of movements across different robot platforms. It shows the control loop can actually handle that kind of coupling.

Taro: I’m still curious how this holds up when we introduce much messier environments, because they flag that as a limitation—the model doesn't explicitly identify low-level dynamics, so it relies on the state prediction being good enough.

Rosa: That’s the caveat: it doesn't try to solve the physics engine itself; it just uses those predicted states as jointly generated variables connecting high-level actions to their physical and visual consequences.

Dev: So, for an engineer, that means you’re not relying on a perfect low-level controller being there; you’re relying on this AI model to implicitly learn the dynamics through the training paths.

Taro: It changes how we think about autonomy, because instead of just trying to correct errors after they happen, this framework builds a more grounded understanding of the interaction from the start.

Rosa: And for anyone listening who only cares about real-world applications, it means AI can start to handle more complex physical tasks on humanoid platforms than before.

Dev: We've got a lot of work left on how to keep this joint target stable and fast enough for deployment in real-time systems.

Taro: Next time we look at papers, I want to see if these dynamics models can generalize across different physical body types, not just the ones they tested.

Yan Yang, *Equal contribution., *Corresponding author., Jikun Rong, *Equal contribution., Minzhao Zhu, *Equal contribution., Zheyi Zhao, *Equal contribution., Qirui Hu, *Equal contribution., Zihan Lan, *Equal contribution., Weixin Mao, *Equal contribution., Yinhao Li

LimX Dynamics

cs.RO, cs.AI

Submitted: 2026-10-08

Updated: 2026-10-08

Project page: https://hwam.vercel.app

The gist: The gist The HWAM introduces a Humanoid World Action Model with joint state–action generation, which makes the robot’s post-execution proprioceptive state an explicit prediction target to address

Key concepts

Action–Execution Gap
This is the discrepancy between the intended movement (the reference action from a high-level policy) and the actual physical motion realized by the robot's joints. The paper addresses this gap by making the robot's resulting state an explicit target for action generation.
Joint Prediction Target (Yt)
This is a combined prediction target that includes both the current joint state (St) and the proposed action (At). It serves as the core input for a shared generator, allowing it to access information about the robot's physical configuration and planned movement simultaneously during training.
Complementary Training Paths
HWAM uses three distinct training methods: Policy path, Forward Dynamics Modeling (FDM), and Inverse Dynamics Modeling (IDM). These paths work together to supervise the model by predicting future visual observations based on actions and resulting states, ensuring the learned actions are physically plausible.
Path-Dependent Masked Cross-Attention
This mechanism connects different training paths—specifically between a video expert and a state-action expert. It allows the model to use information from various modalities (visual transitions, joint trajectories) in a context-aware manner, guiding the learning process across multiple objectives.

Terminology

Summary

The gist The HWAM introduces a Humanoid World Action Model with joint state–action generation, which makes the robot’s post-execution proprioceptive state an explicit prediction target to address the action–execution gap in hierarchical humanoid systems.

Motivation and Problem

Humanoid robots require both task-level motion decisions and execution-level whole-body dynamics control, leading to an action–execution gap where the reference produced by a policy can differ from the motion realized by the robot <ref:2610.12026#pg5> The discrepancy between a high-level reference action and its physical realization complicates action learning through future visual prediction because if this intermediate state is not modeled, the model must learn how the policy action is physically realized and how the realized motion changes the scene <ref:2610.12026#pg6> Dataset statistics reveal larger discrepancies between commanded joint references and measured proprioceptive states in mobile humanoid manipulation than in stationary bimanual manipulation <ref:2610.12026#pg10>.

HWAM Architecture and Joint Target

HWAM addresses this problem by explicitly placing the post-execution state in the action-generation target and then connecting the joint state–action trajectory to visual transitions in both directions <ref:2610.12026#pg4>. The method defines a joint prediction target as Yt = Concat(St, At) ∈ R H×(ds+da), where Yt[h] = [st+h+1; at+h], h = 0,..., H − 1 <ref:2610.12026#pg5>. This joint target is generated by a shared state–action generator, allowing the action channels to access the current partially denoised state channels throughout denoising <ref:2610.12026#pg6>.

Complementary Training Paths

HWAM incorporates post-execution states into three complementary conditional paths:

  1. Policy path: This path jointly denoises state–action trajectories conditioned only on current observations, matching deployment conditions <ref:2610.12026#pg4>.

  2. Forward Dynamics Modeling (FDM): FDM predicts future visual observations conditioned on actions and post-execution states, explicitly accounting for realized body motion when predicting scene changes <ref:2610.12026#pg4>.

  3. Inverse Dynamics Modeling (IDM): IDM reconstructs the joint trajectory from visual transitions, adding supervision in the reverse direction by jointly predicting actions and the body states they produce <ref:2610.12026#pg4>.

Bidirectional Grounding Objectives

The three paths are connected through path-dependent masked cross-attention between a video expert and a state–action expert <ref:2610.12026#pg6>. The training utilizes conditional flow matching objectives for both generated modalities, applying elementwise mean squared error (MSE) to the targets, with separate noise levels τV and τSA for future-video and joint trajectory targets respectively <ref:2610.12026#pg6>. The total loss is optimized as LHWAM = Em∼π [λmLm] = 1/3 (LF + LI + LP) <ref:2610.12026#pg6>.

Evaluation and Results

HWAM achieves the highest success rate among evaluated baselines on three real-robot tasks on the LimX OLI humanoid, achieving a 70.6% success rate on Candy Picking compared to 43.3% for Fast-WAM <ref:2610.12026#pg6>. In controlled ablation on mobile manipulation, the effect of the joint target depends on the visual training recipe, showing that adding state–action prediction increases success under FDM/IDM/Policy training from 55.0% to 73.3%, whereas it decreases success under VGM/Policy training from 60.0% to 45.5% <ref:2610.12026#pg6>. HWAM also achieves the highest reported Clothes Folding task-completion score among evaluated methods on the ALOHA bimanual platform, reaching 86.00 <ref:2610.12026#pg5>. The model also achieves a lowest action MAE (3.816◦) in open-loop prediction diagnostics compared to 4.024◦ for action-only FDM/IDM/Policy <ref:2610.12026#pg6>.

Conclusion

HWAM introduces joint state–action generation, which is motivated by a measurable action–execution gap where low-level control and dynamics cause realized body motion to differ from policy references <ref:2610.12026#pg6>. The method connects execution states and action references to visual transitions through FDM and IDM, alongside a Policy path that generates only state–action trajectories conditioned on current observations <ref:2610.12026#pg6>. HWAM achieves the highest observed success rate on the three evaluated OLI tasks <ref:2610.12026#pg6>. The ablation on mobile manipulation shows that adding the joint target improves task success under FDM/IDM/Policy training and reduces it under VGM/Policy, supporting the combined design in this setting <ref:2610.12026#pg6>. Additional ALOHA results demonstrate its applicability to bimanual deformable-object manipulation beyond the humanoid platform <ref:2610.12026#pg5>. The method does not explicitly identify low-level dynamics, and its predicted states are not sent to the controller; they serve as jointly generated variables that connect high-level actions to their physical and visual consequences <ref:2610.12026#pg6>.

Improvements for AI systems

  1. textbf Realized State Supervision via Joint Target Generation: The system can learn from execution failures by using a joint state–action target defined as Yt = Concat(St, At) ∈ R H×(ds+da) to jointly denoise state and action channels. This allows the model to learn how the reference produced by the policy can differ from the motion realized by the robot.

  2. textbf Bidirectional Grounding for Visual Prediction: The system can connect high-level references to visual outcomes in both directions, as demonstrated by having Forward Dynamics Modeling (FDM) predict future visual observations conditioned on actions and post-execution states and Inverse Dynamics Modeling (IDM) reconstructs the joint trajectory from visual transitions.

  3. textbf Deployment-Matched Policy Learning: The high-level policy can be trained to match deployment conditions because the Policy path jointly denoises state–action trajectories conditioned only on current observations, matching deployment conditions. This ensures that learned references are relevant to the robot's actual capabilities.

  4. textbf Robust Manipulation in Complex Environments: The system achieves superior performance in tasks like Candy Picking (70.6% success rate) and Clothes Folding (86.00 score), indicating it can handle mobile object handling with whole-body coordination and bimanual deformable-object manipulation more effectively than previous baselines.

  5. textbf Adaptive Task Success Under Specific Training Recipes: The system's performance is highly dependent on the training recipe, as shown by the observation that success improves from 55.0% to 73.3% when switching from Action + FDM/IDM/Policy to HWAM Joint state–action FDM/IDM/Policy training in mobile manipulation tasks.

Sources

Related papers