RPL: Learning Robust Humanoid Perceptive Locomotion on Challenging Terrains
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "RPL: Learning Robust Humanoid Perceptive Locomotion on Challenging Terrains".
Dev: Humanoid perceptive locomotion has made significant progress, but achieving robust multi-directional locomotion on complex terrains remains underexplored.
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: So we're looking at RPL: Learning Robust Humanoid Perceptive Locomotion on Challenging Terrains, which tackles the challenge of making humanoids move robustly across tricky ground while carrying loads. The core idea seems to be a two-stage training framework that first trains specialists for different terrains and then combines those skills into one unified policy using multiple depth cameras.
Dev: That makes sense, Rosa; the focus on multi-directional movement and payload robustness is definitely what people in control engineering are interested in, especially regarding stability under real-world conditions where latency and noise are factors.
Taro: I'm curious about how this framework handles unexpected situations; when the world misbehaves, does this system have a mechanism for reacting dynamically beyond just following the learned policies?
Rosa: That’s a fair question, Taro; the paper suggests that Stage one focuses on mastering decoupled locomotion and manipulation skills across various terrains like slopes and stairs, which builds a strong foundation before distillation <ref:2602.03002#pg0,decoupled locomotion and manipulation skills across>.
Dev: And Stage two then takes those experts and distills them into one transformer policy using multi-view depth inputs to achieve robust locomotion across those different environments <ref:2602.03002#pg0>.
Taro: That sounds like a way to handle different terrain types because the system learns specific skills first, rather than trying to learn everything at once from scratch, which is smart when dealing with complex environments.
Rosa: Exactly; and they introduce two specific techniques during distillation—Depth Feature Scaling based on Velocity commands and Random Side Masking—to help the unified transformer policy handle those asymmetric visual inputs better.
Dev: I see the importance of those scaling and masking techniques; that suggests they are explicitly trying to mitigate distribution shift when moving between different view angles or terrain types, which directly relates to loop rate stability.
Taro: If we're talking about generalization, how does Random Side Masking specifically help the system handle unseen terrain widths, as the paper mentions?
Rosa: The Random Side Masking technique randomly masks lateral regions of the depth images to improve generalization to unseen terrain widths by sampling a mode like none, small, or large mask for each environment.
Dev: That’s an adaptive way to handle unknown geometry; it's not just one fixed approach, which is something I appreciate when dealing with unpredictable sensor inputs on the ground.
Paper summary: Taro: And what about the training itself? Stage one trains these terrain-specific experts using privileged height map observations to master those decoupled skills across slopes, stairs up and down, and stepping stones <ref:2602.03002#pg0,privileged height map observations to master>.
Rosa: That’s right; they are using privileged height map observations in Stage one to master those decoupled locomotion and manipulation skills across four distinct terrain families: slopes up to thirty-seven degrees, stairs with different step lengths like twenty-two cm, twenty-five cm, or thirty cm, and stepping stones with gaps.
Dev: From a control standpoint, having separate experts for those specific terrains means the system isn't trying to solve the entire problem simultaneously on every single input; it’s decomposing the complexity first.
Taro: Decomposing the problem sounds effective for autonomy; if one part of the locomotion fails, you might still have some learned skill from a specialized expert that helps maintain stability.
Rosa: And Stage two then takes those terrain-specialized experts and distills them into a single unified multi-view, depth-based transformer policy to enable robust bidirectional locomotion using front and back depth observations <ref:2602.03002#pg0>.
Dev: That distillation step is critical because it merges those specialized knowledge pieces into one coherent system that uses multiple views for better perception, which should help stabilize the overall control loop.
Taro: So the ultimate goal here seems to be moving from highly specialized, decoupled skills to a single general visual policy capable of robust multi-directional locomotion across complex scenarios.
Rosa: That’s the gist of RPL: moving from terrain-specific experts trained on height maps to a unified transformer policy that uses multiple depth cameras for robust movement.
Dev: I'm also interested in the efficiency side; they developed an efficient multi-depth rendering system that achieves a five times speedup over existing pipelines while modeling realistic sensor latency and noise.
Taro: Modeling those realistic sensor imperfections in the rendering pipeline is important because it means the learned policy isn't just trained on perfect data but on data that resembles what a real robot would experience during operation.
Rosa: It’s about making sure the simulation environment closely mimics real-world sensor behavior so that when we deploy this, we don't run into problems because the training was too clean.
Paper summary: Dev: The system achieving that speedup while incorporating latency and noise is a huge win for practical application; it speaks directly to making these kinds of complex learning systems viable on actual hardware.
Taro: If you can train a policy this robustly in simulation with realistic noise models, the potential for real-world deployment on genuinely challenging terrains becomes much more plausible.
Rosa: And the validation shows that this works in the real world too; they demonstrated robust multi-directional locomotion with a two kilogram payload across those varied terrains, including twenty-degree slopes and stepping stones separated by sixty cm gaps <ref:2602.03002#pg0>.
Dev: The fact that it maintains performance with a two kilogram load is significant because it means the control loop can handle the added inertia and dynamic changes without immediately failing, which addresses one of the main concerns in locomotion research.
Taro: That payload robustness is key because real-world tasks aren't just about walking; they involve carrying things while navigating obstacles, which is a much harder problem to solve autonomously.
Rosa: And the ablation studies confirm that both Depth Feature Scaling based on Velocity commands and Random Side Masking are critical; removing them causes the success rate to drop from ten out of ten down to as low as zero out of ten under asymmetric visual inputs or unseen terrain widths.
Dev: That confirms the necessity of those techniques; it shows they aren't just added for show, but are functionally required to maintain robustness when things get messy.
Taro: It shows that the learned policy isn't just lucky; it has learned specific ways to adapt its perception to handle visual ambiguities and geometric variations in the environment.
Rosa: The final deployed controller combines this distilled visual locomotion policy with the blind upper-body policy for whole-body action tracking, which is how they manage the full robot movement.
Dev: Combining that refined locomotion policy with a separate upper-body policy means you’re separating the concerns of walking versus manipulating an object, allowing both parts to be trained effectively within their respective frameworks.
Taro: So the implication is that this two-stage approach allows for modularity in learning; you can specialize skills first and then generalize them into a robust system.
Rosa: Precisely, it offers a pathway to tackling multi-directional locomotion on complex terrains that was previously underexplored because existing methods often rely on simpler assumptions.
Paper summary: Dev: Thinking about the deployment, how long do you expect this system to stay stable in the field before significant degradation occurs?
Taro: That depends heavily on the unseen dynamics of those environments, but if it can generalize well to novel curvature and lighting, its operational lifespan could be quite extended for a wide variety of real-world settings.
Rosa: The authors mention that it generalizes zero-shot to an in-the-wild curved building staircase with thirty cm steps featuring unseen curvature and lighting conditions, which is a strong indicator of its potential outside the controlled lab setting <ref:2602.03002#pg0>.
Dev: That zero-shot capability on unknown visual properties is what really puts this method ahead; it suggests the learned features are more fundamental than just memorizing training data points for specific terrains.
Taro: If that generalization holds up when confronted with unpredictable, dynamic real-world interactions, then the impact on autonomous navigation in unstructured environments could be substantial.
Rosa: The RPL paper provides a solid framework for how to approach multi-directional locomotion by breaking it down into specialized training and unified distillation, which is something we can share with the field roboticists.
Dev: And from an engineering standpoint, the efficiency gains in the rendering system are a tangible improvement that makes running these kinds of complex visual learners much more practical for real-time control loops.
Taro: The overall implication is that by carefully structuring the learning process this way, we can build humanoid systems capable of navigating environments far more diverse and dynamic than what's currently possible.
Rosa: So, to wrap up this discussion on RPL: it’s a two-stage training framework that uses specialized experts distilled into a unified transformer policy with specific techniques like DFSV and RSM for better robustness on challenging terrains with payloads.
Dev: And the real-world validation shows it achieves a six out of ten whole-course success rate on the Unitree G1 humanoid under demanding conditions, which is a solid performance metric.
Taro: The future work mentioned suggests exploring sideways locomotion on discrete terrains like stepping stones and addressing active viewpoint selection for highly occluded scenarios, which points toward where the system can go next.
Rosa: That leaves us wondering how long these systems will need to stay deployed before they face limitations in truly ambiguous situations, but the current results suggest a strong foundation for future development.
Conclusion: Rosa: So, we're wrapping up our discussion on RPL, which is about learning robust humanoid locomotion on tricky terrain using this two-stage training framework.
Dev: Yeah, and I think the authors really nailed how they addressed those real-world issues with latency and failure modes by focusing on the distillation process.
Taro: From an autonomy standpoint, I'm really interested in how much of that robustness comes from those specific techniques like Depth Feature Scaling and Random Side Masking when the environment is truly unexpected.
Rosa: Exactly, because they showed that without those components, the system's success rate dropped drastically when facing asymmetric visual inputs or unknown terrain widths.
Dev: That dependence on those specific distillation techniques tells us a lot about what makes a control loop stable under pressure; it's not just about having a big model, it’s about how that model adapts its perception based on the velocity command.
Taro: And the fact that they validated this with real-world data—carrying a two-kilogram payload across slopes and stairs—shows that these concepts aren't just theoretical exercises; they actually translate to handling physical dynamics.
Rosa: That payload robustness is what makes this work so compelling for field robotics, showing it can manage the inertia of carrying weight while navigating obstacles.
Dev: It’s a significant step toward making these systems practical because it shows the control structure can maintain stability even when things get physically demanding.
Taro: I'm still curious about the real-world duration; how long do you think this kind of learned policy will stay reliable before it starts degrading in a field setting?
Rosa: That’s a great question for our listeners, because while it generalizes well to unseen conditions like curved staircases, we haven't seen long-term degradation data yet.
Dev: I agree; the longevity depends on how well the system handles those long sequences of complex interactions without accumulating errors in its state estimation.
Taro: So, as we look ahead, what are the immediate next steps for this research team to push these capabilities further?
Amazon FAR Co-Lead
cs.RO
Submitted: 2026-02-03
Updated: 2026-10-04
Project page: https://rpl-humanoid.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 87/100
The gist: Humanoid perceptive locomotion has made significant progress, but achieving robust multi-directional locomotion on complex terrains remains underexplored.
Key concepts
- Terrain-specific expert policies
- These are specialized AI models trained in the first stage of RPL. They focus on mastering specific movement skills like climbing slopes or navigating stairs. They use privileged height map data to learn decoupled locomotion and manipulation tasks, acting as specialists for different types of challenging ground.
- Depth Feature Scaling based on Velocity commands (DFSV)
- DFSV is a technique used during policy distillation to make the robot's perception features more stable when moving quickly or facing complex, asymmetric views. It adaptively adjusts the visual input features based on the robot's current velocity command, which helps reduce errors caused by shifting visual distributions.
- Random Side Masking (RSM)
- RSM is a method applied to depth images to improve how well the robot handles terrain it has never seen before. It randomly hides parts of the side of the depth camera view. By sampling different masking modes, the policy learns to generalize better to unseen terrain widths and shapes.
Terminology
Summary
Humanoid perceptive locomotion has made significant progress, but achieving robust multi-directional locomotion on complex terrains remains underexplored. RPL proposes a two-stage training framework that enables multi-directional locomotion on challenging terrains and maintains robustness with payloads.
The gist
RPL first trains terrain-specific expert policies with privileged height map observations to master decoupled locomotion and manipulation skills, which are then distilled into a unified transformer policy leveraging multiple depth cameras for robust multi-directional locomotion.
Framework Overview
RPL is structured as a two-stage framework (Figure 2). Stage 1 involves training terrain-specific expert policies
using privileged height map observations
to master decoupled locomotion and manipulation skills.
These experts are then distilled into a single unified policy in Stage 2, which takes multi-view depth inputs to enable robust multi-directional locomotion. The dual-agent formulation builds on FALCON, where the lower-body policy is conditioned on locomotion goals
and the upper-body policy is conditioned on upper-body joint targets.
Key Distillation Techniques
During distillation into the unified transformer policy, two techniques are introduced to robustify multidirectional locomotion:
-
Depth Feature Scaling based on Velocity commands (DFSV): This adaptively modulates perception features based on the velocity command to
reduce distribution shift under asymmetric multi-view inputs.
The scales are applied element-wise to CNN feature vectors before concatenation, requiringno additional learned parameters.
-
Random Side Masking (RSM): This technique randomly masks lateral regions of depth images to
improve generalization to unseen terrain widths.
For each environment, a mode is sampled from a categorical distribution over modes likenone, small, large,
where larger masks emulate narrower effective widths.
Efficient Multi-Depth System
For scalable depth distillation, an efficient multi-depth rendering system is developed that performs ray-casting against both dynamic robot meshes and static terrain meshes in massively parallel environments. This system achieves a 5× speedup over the depth rendering pipelines in existing simulators while modeling realistic sensor latency, noise, and dropout.
The kernel queries all robot body meshes locally to find an early termination upper bound
before querying the shared terrain mesh.
Terrain Specialization and Reward Design
Stage 1 trains specialized experts for four terrain families: (1) Slopes (inclinations up to 37◦), (2) Stairs Up / Down, and (3) Stepping Stones. The reward design incorporates three terms for stable traversal:
(i) Foot Edge Penalty: Penalizes contacts at terrain edges using precomputed dilated edge masks.
(ii) Foothold Penalty: Samples a grid across the foot sole to penalize invalid coverage.
(iii) Torso Orientation Tracking: Coordinates waist and hip joints for a wide stance-mode tracking range.
Real-World Validation and Robustness
Extensive real-world experiments demonstrate robust multi-directional locomotion with payloads (2 kg) across challenging terrains, including 20◦ slopes, staircases with different step lengths (22 cm, 25 cm, 30 cm), and 25 cm×25 cm stepping stones separated by 60 cm gaps.
Ablation studies confirm that DFSV and RSM are critical; removing them causes a drop in success rates from 10/10 to as low as 0/10
under asymmetric visual inputs or unseen terrain widths. The final deployed controller combines the distilled visual locomotion policy with the blind upper-body policy for whole-body action tracking.
Network Architecture Comparison
The distillation process compares four backbones: (i) CNN+MLP, (ii) CNN+RNN, (iii) CNN+Transformer (Ours), and (iv) U-Net+Transformer. The CNN+Transformer
architecture achieves the lowest distillation loss and the highest deployment success
thanks to its attention-based multi-view fusion. Furthermore, comparing distillation objectives shows that DAgger Only reaches a substantially higher terrain level and converges faster than Hybrid DAgger + RL,
as it preserves higher pairwise cosine similarity between terrains during early training.
Real-World Performance
The method achieves an overall 6/10 whole-course success rate
on the Unitree G1 humanoid for back-and-forth traversal across a 50 m course, including a 2 kg payload. The success is maintained across forward or backward directions and generalizes to zero-shot to an in-the-wild curved building staircase (30 cm step) with unseen curvature and lighting.
Failure modes are primarily concentrated in foot landings near terrain edges
and payload swing.
Limitations
The paper notes two main limitations: first, it does not demonstrate real-world sideways locomotion on discrete terrains like stepping stones; second, DFSV improves robustness but does not explicitly learn active viewpoint selection for highly occluded or ambiguous loco-manipulation scenarios.
Improvements for AI systems
Based on the provided scientific paper, here are specific improvements that can be made to existing AI systems, categorized by the core components of RPL:
)1. Improvement in Locomotion Policy Architecture (Stage 2):
Existing end-to-end policies often fail under asymmetric multi-view inputs or unseen terrain widths. By implementing the RPL architecture—distilling terrain-specific experts into a unified Transformer policy—the AI system can achieve:
The ability to perform robust, long-horizon, multi-directional locomotion (bidirectional) across highly complex and novel terrains (e.g., stairs with varying step lengths or stepping stones with large gaps) while maintaining high stability even when the visual input is asymmetric (i.e., different cameras observing distinct terrain types).
)2. Improvement in Perception System Efficiency (Multi-Depth Rendering):
The current bottleneck in vision-based locomotion is the computational cost of rendering multiple depth views, especially with dynamic meshes. By adopting the RPL multi-depth rendering system:
The AI system can process and fuse real-time depth information from a wide field of view (multiple cameras) against both static terrain and dynamic robot geometry with a significant speedup (reported 5x over existing simulators), enabling faster policy inference required for high-frequency whole-body control.
)3. Improvement in Robustness to Visual Ambiguity (DFSV & RSM):
To handle the inherent noise, occlusion, and lack of training data for unseen terrain widths:
The AI system can maintain stable locomotion when presented with asymmetric visual inputs (e.g., a front camera seeing stairs while the rear sees stepping stones) by adaptively scaling depth features based on commanded velocity (DFSV) and by ignoring irrelevant peripheral vision regions (RSM), leading to near-perfect generalization to unseen terrain widths.
)4. Improvement in Training Stability and Generalization:
The two-stage training framework, leveraging privileged height maps in Stage 1 followed by distillation, allows for better skill isolation:
The AI system can decouple complex skills like locomotion from manipulation during initial learning (Stage 1), leading to more stable expert policies that are then distilled into a unified policy without catastrophic forgetting or instability typical of end-to-end learning on highly multimodal data.
)5. Improvement in Real-World Deployment Capability:
The combination of the optimized policy and robust perception allows for reliable deployment:
The AI system can perform long-horizon, whole-body locomotion under significant external disturbances (e.g., carrying a 2 kg payload) on real hardware, successfully navigating complex courses that combine slopes, stairs of different lengths, and stepping stones with gaps in both forward and backward directions.
Sources
- FALCON: Learning Force-Adaptive Humanoid Loco-Manipulation
- HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit
- ExBody2: Advanced Expressive Humanoid Whole-Body Control
- Expressive Whole-Body Control for Humanoid Robots
- Bridging the Sim-to-Real Gap for Athletic Loco-Manipulation
- OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning
- ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills
- Dynamic Loco-manipulation on HECTOR: Humanoid for Enhanced ConTrol and Open-source Research
- OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
- BeamDojo: Learning Agile Humanoid Locomotion on Sparse Footholds
- Humanoid Parkour Learning
- Gallant: Voxel Grid-based Humanoid Locomotion and Local-navigation across 3D Constrained Terrains
- DPL: Depth-only Perceptive Humanoid Locomotion via Realistic Depth Synthesis and Cross-Attention Terrain Reconstruction
- Learning Perceptive Humanoid Locomotion over Challenging Terrain
- Omni-Perception: Omnidirectional Collision Avoidance for Legged Locomotion in Dynamic Environments
- Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal Transformers
- Hiking in the Wild: A Scalable Perceptive Parkour Framework for Humanoids
- Gait-Adaptive Perceptive Humanoid Locomotion with Real-Time Under-Base Terrain Reconstruction
- MoRE: Mixture of Residual Experts for Humanoid Lifelike Gaits Learning on Complex Terrains
- JAEGER: Dual-Level Humanoid Whole-Body Controller
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving