AdaptManip: Learning Adaptive Whole-Body Object Lifting and Delivery with Online Recurrent State Estimation
summary
The gist
AdaptManip presents a fully autonomous framework for humanoid robots to perform integrated navigation, object lifting, and delivery by training a robust loco-manipulation policy via reinforcement
In short
AdaptManip is a learning framework for humanoid robots to autonomously navigate, lift, and deliver objects without human help. It uses a three-stage strategy involving locomotion, manipulation control, and an online recurrent estimator that tracks the object's position using vision and robot feedback. This allows the robot to handle complex tasks robustly in real-world scenarios.
Key concepts
- Recurrent Object State Estimator
- This module is designed to track the manipulated object's 6D pose in real-time, even when the robot has limited vision or experiences occlusions. It fuses data from visual sensors (like cameras), LiDAR, and internal robot signals (proprioception) using a V-LSTM and MLP structure. This allows the system to maintain a continuous belief about where the object is located during manipulation.
- Whole-Body Base Policy
- This is the foundational reinforcement learning policy trained to generate stable bipedal locomotion for the humanoid robot. It learns how to walk and move its base effectively, providing a stable platform for subsequent manipulation tasks. Its training focuses on tracking base velocity and maintaining joint stability under various physical conditions.
- Manipulation Residual Policy
- This policy is trained on top of the locomotion policy to learn task-specific actions needed for grasping and lifting objects. It takes the estimated object pose as input, allowing it to adapt its movements while ensuring that the overall robot motion remains stable. This component focuses specifically on how to interact with and move the target object.
- Sim-to-Real Transfer
- This refers to the ability of a model trained in a simulated environment (like IsaacLab) to perform well when deployed on physical hardware. AdaptManip demonstrates strong sim-to-real transfer by successfully performing tasks in MuJoCo and achieving real-world success, proving the learned policies generalize effectively from simulation to physical robots.
Terminology used across episodes
This episode discusses
- AdaptManip: Learning Adaptive Whole-Body Object Lifting and Delivery with Online Recurrent State Estimation · Paper Radio
- VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation
- ResMimic: From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning
- OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
- GMT: General Motion Tracking for Humanoid Whole-Body Control
- TWIST: Teleoperated Whole-Body Imitation System
- FALCON: Learning Force-Adaptive Humanoid Loco-Manipulation
- HDMI: Learning Interactive Humanoid Whole-Body Control from Human Videos
- PhysHSI: Towards a Real-World Generalizable and Natural Humanoid-Scene Interaction System
- Whole-Body Bilateral Teleoperation with Multi-Stage Object Parameter Estimation for Wheeled Humanoid Locomanipulation
- Highly Dynamic Quadruped Locomotion via Whole-Body Impulse Control and Model Predictive Control
- Sim-to-Real: Learning Agile Locomotion For Quadruped Robots
- RMA: Rapid Motor Adaptation for Legged Robots
- Track Any Motions under Any Disturbances
- Switch-JustDance: Benchmarking Whole Body Motion Tracking Controllers Using a Commercial Console Game
- CLONE: Closed-Loop Whole-Body Humanoid Teleoperation for Long-Horizon Tasks
The paper
AdaptManip: Learning Adaptive Whole-Body Object Lifting and Delivery with Online Recurrent State Estimation · Read on arXiv
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "AdaptManip: Learning Adaptive Whole-Body Object Lifting and Delivery with Online Recurrent State Estimation".
Dev: AdaptManip presents a fully autonomous framework for humanoid robots to perform integrated navigation, object lifting,
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So we're diving into AdaptManip today. We're looking at a paper that claims to let humanoid robots do navigation, picking up objects, and delivering them all by just using reinforcement learning without any human demonstrations or teleoperation data. It sounds incredibly ambitious for field robotics.
Dev: That’s what the title says: "AdaptManip: Learning Adaptive Whole-Body Object Lifting and Delivery with Online Recurrent State Estimation." The authors are Morgan Byrd, Donghoon Baek, Kartik Garg, Hyunyoung Jung, Daesol Cho, Maks Sorokin, Robert Wright and Sehoon Ha. It’s interesting how they tackle the problem of complex manipulation without relying on pre-recorded human actions.
Taro: I'm curious about the core mechanism here; how can a system learn such intricate tasks autonomously when it hasn't been explicitly shown what to do? We need to understand the architecture behind this framework.
Rosa: Exactly, Taro, that’s the million-dollar question for field robotics. The paper outlines a three-stage strategy: navigation, grasping and lifting, and carrying to the destination. It suggests they tackle these stages sequentially using their onboard sensing capabilities.
Dev: And what really stands out is that they use a recurrent object state estimator alongside a whole-body base policy for locomotion, which is then augmented by a residual manipulation control policy for the actual lifting part. That structure seems designed to keep things stable while learning the task-specific actions.
Taro: The concept of an online, recurrent object state estimator that relies on fusing vision, LiDAR, and proprioception to track the object pose in real time under partial visibility is quite sophisticated. How robust is that estimation when the robot is moving around?
Rosa: That’s a big concern for me from a field perspective. The authors acknowledge that vision-based approaches often struggle with occlusions or limited view frustums during whole-body motion, which is why they moved toward relying exclusively on fully onboard sensing to recurrently estimate the object state, robot–object contact forces, and the robot’s global pose.
Dev: From an engineering standpoint, that real-time tracking capability is crucial for maintaining control loops. The paper mentions that the LiDAR-based robot global position estimator provides drift-robust localization, which is necessary to ensure the robot doesn't get lost while performing navigation and manipulation tasks.
Taro: And when things go wrong, what happens? I want to know what this system does when the world misbehaves during carrying or lifting; it has to be recovery-capable.
Rosa: The methodology is designed specifically for recovery; they decomposed the whole-body loco-manipulation task into stages where each stage addresses distinct requirements, including how the robot navigates, then how it grasps and lifts, and finally how it transports the object to its target.
Title and authors: Dev: Regarding failure modes, I see them focusing on a reward design that heavily weights locomotion stability—things like tracking command inputs for base linear velocity and penalizing joint accelerations—to ensure the robot stays mobile during manipulation.
Taro: That reward structure sounds important because it directly influences how the learned policy adapts to physical constraints, like maintaining balance while under changing contact conditions during transport.
Rosa: The authors augment the manipulation policy's reward with specific objectives like kinematic tracking and slip avoidance terms, which include penalties for excessive relative motion between the robot and the box and discouraging tangential hand motion that indicates slipping.
Dev: That level of detail in the reward function suggests they are very focused on making sure the learned policy isn't just successful in a perfect simulation but can handle real-world physical interactions where slip is a constant threat.
Taro: Considering all this, what are the key improvements they propose over prior work, beyond just the combination of components? What makes this approach fundamentally different from existing imitation learning methods?
Rosa: The main improvement seems to be training the entire framework—locomotion and manipulation—via reinforcement learning without any human demonstrations or teleoperation data at all. This is a significant step away from brittle imitation learning approaches that usually fail when faced with disturbances.
Dev: The paper highlights this in Table I, which shows their method achieves success across several metrics, including the absence of human demonstrations and teleoperation data, suggesting a high capability in the 'NoHumRef' and 'NoTeleOp' categories.
Taro: That implies that the system is learning a generalized understanding of how to interact with an object through trial and error guided by the reward signals, rather than just memorizing specific paths shown by a human operator.
Rosa: Precisely, Taro; it’s about training a policy that can adapt to situations it hasn't seen before, relying on the recurrent estimator for continuous pose updates during the task.
Dev: From an engineering loop rate perspective, we have to consider how fast that entire closed-loop system operates. The success hinges on how quickly the object state estimator can provide feedback to update the residual policy for stable lifting and delivery.
Taro: If the estimation latency is too high, even a good locomotion policy might struggle to maintain stability during dynamic contact events. What are their findings on how this latency impacts performance?
Rosa: They demonstrated that the recurrent estimator's advantage is visible in hardware experiments: while visual estimates degrade during floating-base motion, the recurrent estimator continues to track the object by leveraging robot proprioception and policy actions.
Dev: That confirmation of sim-to-real transfer via that recurrent tracking mechanism is a huge deal for deployment; it shows they aren't just getting lucky in simulation, but the underlying estimation logic holds up when the robot is actually moving.
Title and authors: Taro: So, what are the broader implications here for autonomous systems operating in unstructured environments? If we can achieve this level of autonomy without pre-programming every single interaction, it opens up possibilities in many areas.
Rosa: It suggests a path toward robots that are far more flexible and capable of handling the messy reality of real-world tasks, moving beyond highly constrained lab settings.
Dev: I think the implication for control engineers is that we can rely less on painstakingly hand-tuning every single contact force or trajectory for a specific task because the learning framework handles that adaptation through the residual policy.
Taro: And from an autonomy research standpoint, it validates the idea of hierarchical learning where a stable base locomotion policy provides the necessary foundation for higher-level, task-specific manipulation capabilities.
Rosa: So, to wrap up this discussion on AdaptManip: it successfully combined multimodal inputs—LiDAR, vision from camera and AprilTag, and proprioception—to maintain a recurrent belief of the box pose while using hierarchical reinforcement learning to pick up and carry an object.
Dev: The overall result shows that this framework can successfully navigate, lift, and deliver a box in real-world conditions on physical hardware, achieving a seventy-five percent success rate during sim-to-sim transfer to the unseen MuJoCo environment.
Taro: The real world deployment aspect is what really excites me; seeing it successfully operate autonomously on a Unitree G1 humanoid robot using only onboard sensing confirms the viability of this approach in practical applications.
Rosa: It’s certainly a strong demonstration of how structured learning, when combined with robust state estimation, can handle complex locomotion and manipulation in one integrated system.
Dev: We've discussed the technical structure and the performance metrics; it seems like AdaptManip is pushing toward systems that can handle more dynamic contact situations reliably.
Taro: I think this work sets a new benchmark for learning policies for complex, contact-rich interactions without explicit demonstrations, which could influence how we design agents for physical tasks in general.
Rosa: We’ve covered the title, the strategy, the estimation module, and the hardware validation of AdaptManip today. It really shows how integrating these sensing modalities creates a more resilient system for field robotics.
Dev: It’s clear that for control engineers, this reinforces the importance of designing policies that are inherently robust to estimation errors, which is exactly what the recurrent estimator aims to mitigate.
Taro: The ability of this system to maintain an object state belief while walking around demonstrates a level of self-awareness in the robot's perception that is very promising for future autonomous systems.
Rosa: That’s all we have time for today regarding AdaptManip, and it really shows the power of combining these different learning and estimation techniques.
The paper's summary: Rosa: So, to recap, AdaptManip is this framework that lets humanoid robots do navigation, lifting, and delivery autonomously by learning everything through reinforcement learning without needing any human demonstrations or teleoperation data.
Dev: That's the core idea—training a whole-body locomotion policy alongside a manipulation policy that adapts based on real-time object information—which sounds way more flexible than traditional programmed motion sequences.
Taro: I'm really interested in how they handle the "online" part of the state estimation; it’s not just a snapshot, but a continuous update loop that lets the robot recover from things going wrong during the whole process.
Rosa: Exactly, Taro; this recurrent estimator fuses vision, LiDAR odometry, and proprioceptive data to keep track of where the object is even when things get occluded by the robot's body or hands.
Dev: And from a control perspective, that continuous tracking is what allows the residual manipulation policy to adjust its actions on the fly, ensuring stability even if the object's pose estimate has some error.
Taro: It means we're looking at a system that doesn't just fail when it encounters something unexpected; it actively tries to maintain its understanding of the environment and the task objective, which is a big step for autonomy.
Rosa: And the results are pretty compelling, showing that this system can actually work in the real world on physical hardware, like a Unitree G1 robot.
Dev: The success rate they reported during sim-to-sim transfer to unseen environments was seventy-five percent, which is a solid figure, especially when you compare it to other complex learning setups.
Taro: That simulation performance, particularly in that unseen environment test, really speaks to the robustness of their method; it suggests the policy isn't just memorizing simulation data but is actually generalizing its understanding.
Rosa: So what this implies for us is that we could potentially design robots that are far less reliant on painstakingly hand-tuning every single interaction for a specific task, opening the door for much more flexible field operations.
Dev: That flexibility comes at a cost, though; we have to worry about the loop rate of that entire closed-loop system and how quickly those state estimates can feed back into the control actions, especially during dynamic contact events.
Taro: If the estimation latency is too high, even a good locomotion policy might struggle to maintain stability during dynamic contact events, so I'm curious if they found any specific thresholds for acceptable performance.
Rosa: They did suggest that the continuous nature of the recurrent estimator helps mitigate those latency issues by leveraging proprioception and policy actions to keep tracking going even when vision is limited.
Dev: That’s a crucial point; it means the architecture is designed to be inherently fault-tolerant regarding sensory input, which is something we need to build into our next generation of controllers.
Taro: It really shows that the combination of hierarchical RL with a recurrent belief system is a powerful way to tackle complex, contact-rich interactions where failure is inevitable and recovery is key.
The paper's improvements: Taro: So, to sum up the improvements, AdaptManip isn't just about adding components; it’s about building this structured three-stage strategy—navigation, grasping, and carrying—and then training the whole thing end-to-end using reinforcement learning without any human help.
Rosa: That's right; they are not relying on someone showing the robot exactly how to walk or how to lift a box; the AI learns all those complex behaviors through trial and error guided by a carefully designed reward system.
Dev: What I find particularly interesting is how they separate the learning into two distinct policies: a base locomotion policy for stable walking, and then a residual manipulation policy that only focuses on adapting to the object once it's grasped.
Taro: That hierarchical structure seems smart because it lets the robot focus its learning efforts where they matter most, ensuring it maintains fundamental stability while simultaneously learning the fine motor skills of grasping and delivering.
Rosa: And I think the real muscle here is that recurrent object state estimator, which acts like a persistent memory for the object's location, allowing the system to recover even when visual input gets messy or temporarily missing.
Dev: That memory function is critical because it enables that closed-loop recovery you mentioned; without that continuous update on what the box is doing, any momentary slip would probably cause a complete failure of the manipulation task.
Taro: It means the system can actually handle failures in a dynamic way, instead of just stopping or freezing when something goes wrong during carrying.
Rosa: Exactly, and when we look at sim-to-real transfer, they suggest a specific method where they train the policy using noisy or masked ground truth poses to explicitly model those visual estimation errors and occlusions from the start.
Dev: That explicit modeling of uncertainty in training is smart because it teaches the policy how to behave predictably when its sensor data isn't perfect in reality, which is vital for deployment.
Taro: It suggests a new way of approaching sim-to-real transfer where you don't just hope it works; you train the system to be robust against the specific types of errors it will encounter in the real world.
Rosa: So, this framework moves beyond just achieving a single task; it’s building an adaptable robot capable of performing a sequence of complex, multi-step actions autonomously in an unstructured setting.
Dev: The implication for control systems is that we can design controllers that are inherently resilient to estimation noise by integrating recurrent feedback mechanisms directly into the policy's state space, rather than trying to filter it out externally.
Taro: If this approach scales, I imagine we could see agents capable of performing much more intricate tasks in real-world environments where human supervision isn't feasible at all.
Conclusion: Rosa: So we've seen how AdaptManip uses a combination of recurrent state estimation and hierarchical reinforcement learning to let humanoids autonomously navigate, pick up, and deliver objects without any human guidance or pre-recorded demonstrations.
Dev: That’s right; it really shows how combining robust estimation with adaptive control policies can handle the complexity of whole-body manipulation in real environments.
Taro: I think the biggest impact is how this framework addresses the "what if" scenarios; it’s not just about following a path, but about maintaining competence when unexpected disturbances happen during that delivery.
Rosa: It really is; we're looking at a future where robots aren't just programmed for specific routes but are truly capable of adapting their whole strategy in real-time based on what they perceive.
Dev: From an engineering standpoint, the robustness against failure modes is key here; that online estimation loop means the system can handle transient errors and continue its mission instead of crashing immediately.
Taro: And that ability to recover during manipulation suggests we might see agents capable of handling much more intricate tasks in unstructured settings where human supervision simply isn't available.
Rosa: It’s certainly exciting to think about robots that can operate effectively outside the lab for extended periods, navigating and completing complex tasks on their own using only onboard sensing.
Dev: We've seen strong results on physical hardware, which gives us confidence that the simulation-to-real transfer isn't just a lucky fluke but is based on sound mechanisms.
Taro: That validation across different environments shows that the generalization achieved by this recurrent estimator is quite solid for complex locomotion and manipulation.
Rosa: So, to wrap up, AdaptManip provides a powerful learning-based framework for whole-body humanoid loco-manipulation that tackles navigation, lifting, and delivery autonomously using onboard sensing.
Dev: It sets a high bar for how we design policies that prioritize stability while allowing the system to adapt its behavior based on real-time environmental feedback.
Taro: We've seen how this architecture can handle failures gracefully during complex interaction sequences, which is where the true autonomy lies in field robotics.
More episodes
- 2610.11768-Narrow and Deep: An Ontology Tower as the Knowledge of an LLM Agent for an Industrial Equipment System
- 2610.11904-Large-Scale Partition-Based RIS Beamforming For Uplink RIS-Equipped Multi-User Systems: Asymptotic Analysis
- 2610.11885-Redefining fuel poverty: Introducing the temporal equity framework (TEF)
- 2610.11900-Reach-Stabilize Control of Control-Affine Systems with Unknown Affine Parameters
- 2610.11964-From Asymptotic to Designer-Assigned-Time Control: A Review of Stability Notions, Design Mechanisms, and Controller Architectures
- 2610.12226-Stabilization of Unidirectional First-Order PDE-ODE Coupled Systems with Boundary and Distributed Input Delays
- 2610.12028-Policy Synthesis for Finite Populations of MDP Agents under Aggregate Reach-Avoid Chance Constraints
- 2610.12103-Predefined-Time Integral Reinforcement Learning for Unknown Nonlinear Systems via Inverse-Optimal Design
- 2610.12110-Adaptive dynamic programming using Lyapunov function constraints
- 2610.12324-Convex Safety Filtering via Spectral Selection for Nonconvex Safe Sets