DAMP: Humanoid Locomotion via Denoised Belief Learning and Adversarial Motion Priors

arXiv:2610.11505 · cs.RO · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "DAMP: Humanoid Locomotion via Denoised Belief Learning and Adversarial Motion Priors".

Dev: The gist: DAMP introduces a reinforcement learning framework for robust and naturalistic humanoid locomotion over challenging terrains by implicitly inferring privileged and other task-relevant latent information using recurrent neural networks.

Rosa: First, who's behind it and why it matters.

Paper summary: Dev: So wrapping up with DAMP: Humanoid Locomotion via Denoised Belief Learning and Adversarial Motion Priors. The authors developed this framework to let humanoid robots traverse complex terrains without needing perfect prior knowledge of the environment's state.

Rosa: They use recurrent neural networks to infer latent information from observations, which they then align with task objectives, making the policy learning more robust and goal-consistent than standard approaches that rely only on immediate sensor data.

Taro: What does this mean practically for a robot that needs to operate in a messy factory or an uneven outdoor area?

Dev: It means the system can maintain smooth, natural locomotion over challenging surfaces like stairs or slopes, even if the underlying state information is partially missing from what’s currently being observed.

Rosa: The implication is that we might move toward systems that don't need a perfect map of everything to function reliably in unstructured settings.

Taro: It suggests that learning how to handle uncertainty and infer the hidden context is just as important as learning the direct control commands.

Dev: They show this works on the Noetix N2 robot, validating its ability to perform this agile locomotion across stairs and rough terrains during real-world testing.

Rosa: This paper demonstrates that combining belief learning with adversarial priors provides a solid method for achieving reliable humanoid movement in unpredictable environments.

Conclusion: Rosa: So, we've been looking at this paper called DAMP by these authors, and basically, they’re trying to figure out how robots can walk over really rough ground without knowing everything about what’s around them beforehand.

Dev: Right, it uses this reinforcement learning framework that tries to guess the hidden state of the world while also trying to make sure the robot moves in a way that looks natural.

Taro: I'm curious about what they mean by "denoised belief learning"—is it just making educated guesses about things we can't actually see?

Rosa: Well, they use these recurrent neural networks to keep track of time and infer what the robot *should* be seeing, even when the sensors are a bit noisy.

Dev: From an engineering standpoint, I'm checking how fast this whole loop runs; if it’s too slow, all that inference doesn't help you get a stable gait.

Taro: But what happens when the world surprises the robot—like a sudden unexpected step on a rock—does this framework handle that unpredictability?

Rosa: They’ve also added this adversarial motion prior, which basically trains the robot to mimic expert movements, so it learns what "natural" walking looks like.

Dev: And they combine that task objective reward with the style reward from the discriminator, trying to keep the movement both functional and smooth.

Taro: So it’s not just about getting from point A to B; it’s about getting there in a way that feels right for a human walking on uneven terrain.

Rosa: Exactly. The paper suggests this integration of inference and imitation guidance helps create a much more robust system than just standard reinforcement learning setups.

Dev: And the real validation came from testing it on the Noetix N2 robot, showing it actually works well over stairs and slopes in the real world.

Taro: It seems like this approach moves us closer to robots that can navigate messy, unstructured environments without needing perfect pre-programming for every single corner.

Rosa: That’s where we’re heading, so next time we look at a paper like this, we gotta ask how long these systems can actually keep performing reliably outside of a controlled lab setting.

Puying Shen, Wenhao Cui, Huaxing Huang, Bangyu Qin, Shengtao Li, Ziyang Dong, Guoteng Zhang

cs.RO

Submitted: 2026-10-08

Updated: 2026-10-08

The gist: The gist: DAMP introduces a reinforcement learning framework for robust and naturalistic humanoid locomotion over challenging terrains by implicitly inferring privileged and other task-relevant

Key concepts

Denoised World Learning (DWL)
This module uses an encoder-decoder architecture, powered by a recurrent network, to process robot observations. It aims to extract a latent state representing the environment while reconstructing it toward privileged information like ground friction and terrain elevation. This helps the system learn robust representations from available sensory data.
Adversarial Motion Prior (AMP)
AMP uses an adversarial learning setup with a discriminator network to guide the policy toward generating motions similar to expert demonstrations. A style reward derived from this discriminator encourages the robot to produce natural, human-like locomotion, improving overall movement quality.
Proprioceptive-to-Privileged State Inference
The framework infers crucial latent information by mapping raw proprioceptive observations (like joint positions and velocities) to a privileged state. This process implicitly learns important task details that are not directly observed, allowing the policy to make better decisions in complex, unstructured environments.
Proximal Policy Optimization (PPO)
PPO is the core reinforcement learning algorithm used to train the entire end-to-end framework. It optimizes all components—the policy network and value network—simultaneously by ensuring that policy updates are stable and do not drastically change the learned behavior, leading to reliable locomotion.

Terminology

Summary

The gist: DAMP introduces a reinforcement learning framework for robust and naturalistic humanoid locomotion over challenging terrains by implicitly inferring privileged and other task-relevant latent information using recurrent neural networks.

Framework Overview

DAMP is a reinforcement learning framework aimed at achieving robust and naturalistic humanoid locomotion over challenging terrains, with the assumption that no perceived information is available<ref:2610.11505#pg2>. The framework leverages recurrent neural networks to capture temporal dependencies and implicitly infer privileged and other task-relevant latent information<ref:2610.11505#pg4>. This end-to-end framework achieves transfer learning from simulation to real-world environments, demonstrating the proposed method’s robustness and generalization capabilities<ref:2610.11505#pg4>. The DAMP framework is built upon three principal components: denoised world learning (DWL) module, the critic network, and AMP which has been introduced in Section II<ref:2610.11505#pg4>. During training, the DWL module processes proprioceptive observations available at deployment, while the decoder reconstructs the latent variable toward privileged information, encouraging latent alignment with richer supervisory signals<ref:2610.11505#pg4>. The overall framework is trained using Proximal Policy Optimization (PPO), enabling end-to-end optimization of all components within a unified reinforcement learning paradigm<ref:2610.11505#pg4>.

Policy and Value Networks

The policy network consists of a recurrent backbone and a feedforward actor head, specifically an LSTM-based recurrent neural network encodes temporal dependencies from sequential observations, and a single-head multilayer perceptron (MLP) serves as the actor to output joint position commands<ref:2610.11505#pg4>. At each timestep t, the policy takes ot as input at each timestep, where ot = ct = (vx t, vy t, ω yaw t), ωt is the base angular velocity measured by the IMU, gt represents the gravity direction expressed in the body frame, θt and θ˙t denote joint positions and velocities obtained from joint encoders, and at-1 is the action applied at the previous timestep<ref:2610.11505#pg4>. The value network takes a privileged state st as input, where the full critic state is defined as st = ot, vt, mt, µt, ρt, kp, kd, ηt, c foot t, ht, (6)<ref:2610.11505#pg4>. The action at corresponds to the target positions for the robot joints<ref:2610.11505#pg4>. The joint torques τ t are computed using a PD controller, where τ t = kp(at − θt) − kd θ˙t<ref:2610.11505#pg4>.

Adversarial Motion Prior (AMP)

To encourage natural and human-like locomotion, the framework incorporates the Adversarial Motion Prior (AMP) framework, which guides the policy toward generating motions consistent with expert demonstrations through adversarial learning<ref:2610.11505#pg3>. For adversarial training, a discriminator Dψ(s It, s It+1) is introduced to distinguish transitions from expert demonstrations and those generated by the policy<ref:2610.11505#pg3>. The adversarial loss is defined as LAMP = 1/2 E(s It,s It+1)∼DE Dψ(s It, s It+1) − 1/2 2 + 1/2 E(s It,s It+1)∼Dπ Dψ(s It, s It+1) + 1/2 2 + λGP Eˆs∼Dinterp ∇ˆsDψ(ˆs) 2<ref:2610.11505#pg3>. The style reward derived from the discriminator is defined as rstyle(s It, s It+1) = max 0, 1 − 1/4 Dψ(s It, s It+1)<ref:2610.11505#pg3>. The final reward combines the task objective and the adversarial motion prior as rt = α rtask(st, at) + β rstyle(s It, s It+1)<ref:2610.11505#pg3>.

Denoising Model and Constraints

The Denoising World Model (DWL) module uses an encoder-decoder architecture to estimate and process observation data<ref:2610.11505#pg5>. The recurrent encoder extracts the latent state zt, which is then decoded into a vector of the same dimension as the underlying state st<ref:2610.11505#pg5>. The denoising loss is formulated as Ldenoise = ˜st − st 2 2 + λrzt1, where λr regulates the strength of latent regularization<ref:2610.11505#pg5>. Incorporating privileged signals, including ground friction, actuator torques, and terrain elevation, facilitates accurate online adaptation and implicit system identification<ref:2610.11505#pg5>. To address limitations in smooth policy output and stable policy updates for legged robots, a Lipschitz Continuity Penalty (LCP) is employed<ref:2610.11505#pg5>. The gradient penalty is LGP = Es,a∼D ∇s log π(as) 2 2<ref:2610.11505#pg5>.

Training and Validation

The framework was trained using the NVIDIA IsaacGym simulator, completing 7000 iterations of training in 5 hours on an NVIDIA RTX 4090 graphics card<ref:2610.11505#pg5>. Domain randomization is applied to the robot’s simulation parameters and environment properties to enhance robustness and facilitate smooth sim-toreal transfer<ref:2610.11505#pg5>. Simulation experiments show that DAMP (ours) consistently outperforms ablated variants, demonstrating superior robustness<ref:2610.11505#pg6>. Real-world experiments on the Noetix N2 robot demonstrate superior capability across diverse terrains, including stairs, slopes, and discrete rough terrains<ref:2610.11505#pg7>. The results indicate that integrating proprioceptive-to-privileged state inference with reinforcement learning provides a robust and effective pathway toward reliable humanoid locomotion in unstructured environments<ref:2610.11505#pg8>.

Conclusion

Overall, the results indicate that integrating proprioceptive-to-privileged state inference with reinforcement learning provides a robust and effective pathway toward reliable humanoid locomotion in unstructured environments<ref:2610.11505#pg8>. The framework further incorporates LCP-based constraints during policy optimization to improve robustness and stability<ref:2610.11505#pg8>. Simulation ablation studies demonstrate that both historical context for latent state inference and imitation guidance are essential for maintaining stability and task performance<ref:2610.11505#pg8>. Real-world experiments further validate the system’s agile and adaptive locomotion capabilities across stairs, slopes, and discrete terrains<ref:2610.11505#pg8>. The results indicate that integrating proprioceptive-to-privileged state inference with reinforcement learning provides a robust and effective pathway toward reliable humanoid locomotion in unstructured environments<ref:2610.11505#pg8>. The framework further incorporates LCP-based constraints during policy optimization to improve robustness and stability<ref:2610.11505#pg8>. Simulation ablation studies demonstrate that both historical context for latent state inference and imitation guidance are essential for maintaining stability and task performance<ref:2610.11505#pg8>. Real-world experiments further validate the system’s agile and adaptive locomotion capabilities across stairs, slopes, and discrete terrains<ref:2610.11505#pg8>. The results indicate that integrating proprioceptive-to-privileged state inference with reinforcement learning provides a robust and effective pathway toward reliable humanoid locomotion in unstructured environments<ref:2610.11505#pg8>. The framework further incorporates LCP-based constraints during policy optimization to improve robustness and stability<ref:2610.11505#pg8>. Simulation ablation studies demonstrate that both historical context for latent state inference and imitation guidance are essential for maintaining stability and task performance<ref:2610.11505#pg8>. Real-world experiments further validate the system’s agile and adaptive locomotion capabilities across stairs, slopes, and discrete terrains<ref:2610.11505#pg8>. The results indicate that integrating proprioceptive-to-privileged state inference with reinforcement learning provides a robust and effective pathway toward reliable humanoid locomotion in unstructured environments<ref:2610.11505#pg8>.

Improvements for AI systems

  1. The DAMP framework can achieve robust and goal-consistent policy learning by leveraging recurrent neural networks to capture temporal dependencies and implicitly infer privileged and other task-relevant latent information. This allows the system to perform stable and task-compliant humanoid locomotion with enhanced motion naturalness over challenging terrains.

  2. The incorporation of the Adversarial Motion Prior (AMP) guides the policy toward generating motions consistent with expert demonstrations, as defined by the adversarial loss equation (2). This promotes motion diversity and training stability, which is shown to accelerate locomotion training compared to other methods.

  3. The system can be made more robust against noise and task-irrelevant variations by using the denoising loss formulation: Ldenoise = ˜st − st2 + λrzt1, (10). This enforces a latent state denoising objective, ensuring the encoder and actor focus on essential state information rather than noise and task-irrelevant variations.

  4. The policy's output can be made smoother during training by incorporating the Lipschitz Continuity Penalty (LCP): LGP = Es,a∼D ∇s log π(as) squared, which enforces a constraint on the gradient of the policy, thereby reducing abrupt changes in actions across nearby states.

  5. The system can be enhanced for more informative value estimation by utilizing the full critic state defined as: st = ot, vt, mt, µt, ρt, kp, kd, ηt, c foot t, ht, (6), which includes details like ground friction, actuator torques, and terrain height scan. This facilitates a more accurate value estimation for policy optimization.

  6. The robot can be adapted to real-world deployment by retaining only the necessary components: During deployment, only the encoder and actor are retained, while the decoder and privileged supervision are discarded. This results in a hardware-deployable policy that is efficient for practical use.

Sources

Related papers