AdaptManip: Learning Adaptive Whole-Body Object Lifting and Delivery with Online Recurrent State Estimation

arXiv:2602.14363 · cs.RO, cs.LG · Submitted 2026-02-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "AdaptManip: Learning Adaptive Whole-Body Object Lifting and Delivery with Online Recurrent State Estimation".

Dev: AdaptManip presents a fully autonomous framework for humanoid robots to perform integrated navigation, object lifting,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So we're diving into AdaptManip today. We're looking at a paper that claims to let humanoid robots do navigation, picking up objects, and delivering them all by just using reinforcement learning without any human demonstrations or teleoperation data. It sounds incredibly ambitious for field robotics.

Dev: That’s what the title says: "AdaptManip: Learning Adaptive Whole-Body Object Lifting and Delivery with Online Recurrent State Estimation." The authors are Morgan Byrd, Donghoon Baek, Kartik Garg, Hyunyoung Jung, Daesol Cho, Maks Sorokin, Robert Wright and Sehoon Ha. It’s interesting how they tackle the problem of complex manipulation without relying on pre-recorded human actions.

Taro: I'm curious about the core mechanism here; how can a system learn such intricate tasks autonomously when it hasn't been explicitly shown what to do? We need to understand the architecture behind this framework.

Rosa: Exactly, Taro, that’s the million-dollar question for field robotics. The paper outlines a three-stage strategy: navigation, grasping and lifting, and carrying to the destination. It suggests they tackle these stages sequentially using their onboard sensing capabilities.

Dev: And what really stands out is that they use a recurrent object state estimator alongside a whole-body base policy for locomotion, which is then augmented by a residual manipulation control policy for the actual lifting part. That structure seems designed to keep things stable while learning the task-specific actions.

Taro: The concept of an online, recurrent object state estimator that relies on fusing vision, LiDAR, and proprioception to track the object pose in real time under partial visibility is quite sophisticated. How robust is that estimation when the robot is moving around?

Rosa: That’s a big concern for me from a field perspective. The authors acknowledge that vision-based approaches often struggle with occlusions or limited view frustums during whole-body motion, which is why they moved toward relying exclusively on fully onboard sensing to recurrently estimate the object state, robot–object contact forces, and the robot’s global pose.

Dev: From an engineering standpoint, that real-time tracking capability is crucial for maintaining control loops. The paper mentions that the LiDAR-based robot global position estimator provides drift-robust localization, which is necessary to ensure the robot doesn't get lost while performing navigation and manipulation tasks.

Taro: And when things go wrong, what happens? I want to know what this system does when the world misbehaves during carrying or lifting; it has to be recovery-capable.

Rosa: The methodology is designed specifically for recovery; they decomposed the whole-body loco-manipulation task into stages where each stage addresses distinct requirements, including how the robot navigates, then how it grasps and lifts, and finally how it transports the object to its target.

Title and authors: Dev: Regarding failure modes, I see them focusing on a reward design that heavily weights locomotion stability—things like tracking command inputs for base linear velocity and penalizing joint accelerations—to ensure the robot stays mobile during manipulation.

Taro: That reward structure sounds important because it directly influences how the learned policy adapts to physical constraints, like maintaining balance while under changing contact conditions during transport.

Rosa: The authors augment the manipulation policy's reward with specific objectives like kinematic tracking and slip avoidance terms, which include penalties for excessive relative motion between the robot and the box and discouraging tangential hand motion that indicates slipping.

Dev: That level of detail in the reward function suggests they are very focused on making sure the learned policy isn't just successful in a perfect simulation but can handle real-world physical interactions where slip is a constant threat.

Taro: Considering all this, what are the key improvements they propose over prior work, beyond just the combination of components? What makes this approach fundamentally different from existing imitation learning methods?

Rosa: The main improvement seems to be training the entire framework—locomotion and manipulation—via reinforcement learning without any human demonstrations or teleoperation data at all. This is a significant step away from brittle imitation learning approaches that usually fail when faced with disturbances.

Dev: The paper highlights this in Table I, which shows their method achieves success across several metrics, including the absence of human demonstrations and teleoperation data, suggesting a high capability in the 'NoHumRef' and 'NoTeleOp' categories.

Taro: That implies that the system is learning a generalized understanding of how to interact with an object through trial and error guided by the reward signals, rather than just memorizing specific paths shown by a human operator.

Rosa: Precisely, Taro; it’s about training a policy that can adapt to situations it hasn't seen before, relying on the recurrent estimator for continuous pose updates during the task.

Dev: From an engineering loop rate perspective, we have to consider how fast that entire closed-loop system operates. The success hinges on how quickly the object state estimator can provide feedback to update the residual policy for stable lifting and delivery.

Taro: If the estimation latency is too high, even a good locomotion policy might struggle to maintain stability during dynamic contact events. What are their findings on how this latency impacts performance?

Rosa: They demonstrated that the recurrent estimator's advantage is visible in hardware experiments: while visual estimates degrade during floating-base motion, the recurrent estimator continues to track the object by leveraging robot proprioception and policy actions.

Dev: That confirmation of sim-to-real transfer via that recurrent tracking mechanism is a huge deal for deployment; it shows they aren't just getting lucky in simulation, but the underlying estimation logic holds up when the robot is actually moving.

Title and authors: Taro: So, what are the broader implications here for autonomous systems operating in unstructured environments? If we can achieve this level of autonomy without pre-programming every single interaction, it opens up possibilities in many areas.

Rosa: It suggests a path toward robots that are far more flexible and capable of handling the messy reality of real-world tasks, moving beyond highly constrained lab settings.

Dev: I think the implication for control engineers is that we can rely less on painstakingly hand-tuning every single contact force or trajectory for a specific task because the learning framework handles that adaptation through the residual policy.

Taro: And from an autonomy research standpoint, it validates the idea of hierarchical learning where a stable base locomotion policy provides the necessary foundation for higher-level, task-specific manipulation capabilities.

Rosa: So, to wrap up this discussion on AdaptManip: it successfully combined multimodal inputs—LiDAR, vision from camera and AprilTag, and proprioception—to maintain a recurrent belief of the box pose while using hierarchical reinforcement learning to pick up and carry an object.

Dev: The overall result shows that this framework can successfully navigate, lift, and deliver a box in real-world conditions on physical hardware, achieving a seventy-five percent success rate during sim-to-sim transfer to the unseen MuJoCo environment.

Taro: The real world deployment aspect is what really excites me; seeing it successfully operate autonomously on a Unitree G1 humanoid robot using only onboard sensing confirms the viability of this approach in practical applications.

Rosa: It’s certainly a strong demonstration of how structured learning, when combined with robust state estimation, can handle complex locomotion and manipulation in one integrated system.

Dev: We've discussed the technical structure and the performance metrics; it seems like AdaptManip is pushing toward systems that can handle more dynamic contact situations reliably.

Taro: I think this work sets a new benchmark for learning policies for complex, contact-rich interactions without explicit demonstrations, which could influence how we design agents for physical tasks in general.

Rosa: We’ve covered the title, the strategy, the estimation module, and the hardware validation of AdaptManip today. It really shows how integrating these sensing modalities creates a more resilient system for field robotics.

Dev: It’s clear that for control engineers, this reinforces the importance of designing policies that are inherently robust to estimation errors, which is exactly what the recurrent estimator aims to mitigate.

Taro: The ability of this system to maintain an object state belief while walking around demonstrates a level of self-awareness in the robot's perception that is very promising for future autonomous systems.

Rosa: That’s all we have time for today regarding AdaptManip, and it really shows the power of combining these different learning and estimation techniques.

The paper's summary: Rosa: So, to recap, AdaptManip is this framework that lets humanoid robots do navigation, lifting, and delivery autonomously by learning everything through reinforcement learning without needing any human demonstrations or teleoperation data.

Dev: That's the core idea—training a whole-body locomotion policy alongside a manipulation policy that adapts based on real-time object information—which sounds way more flexible than traditional programmed motion sequences.

Taro: I'm really interested in how they handle the "online" part of the state estimation; it’s not just a snapshot, but a continuous update loop that lets the robot recover from things going wrong during the whole process.

Rosa: Exactly, Taro; this recurrent estimator fuses vision, LiDAR odometry, and proprioceptive data to keep track of where the object is even when things get occluded by the robot's body or hands.

Dev: And from a control perspective, that continuous tracking is what allows the residual manipulation policy to adjust its actions on the fly, ensuring stability even if the object's pose estimate has some error.

Taro: It means we're looking at a system that doesn't just fail when it encounters something unexpected; it actively tries to maintain its understanding of the environment and the task objective, which is a big step for autonomy.

Rosa: And the results are pretty compelling, showing that this system can actually work in the real world on physical hardware, like a Unitree G1 robot.

Dev: The success rate they reported during sim-to-sim transfer to unseen environments was seventy-five percent, which is a solid figure, especially when you compare it to other complex learning setups.

Taro: That simulation performance, particularly in that unseen environment test, really speaks to the robustness of their method; it suggests the policy isn't just memorizing simulation data but is actually generalizing its understanding.

Rosa: So what this implies for us is that we could potentially design robots that are far less reliant on painstakingly hand-tuning every single interaction for a specific task, opening the door for much more flexible field operations.

Dev: That flexibility comes at a cost, though; we have to worry about the loop rate of that entire closed-loop system and how quickly those state estimates can feed back into the control actions, especially during dynamic contact events.

Taro: If the estimation latency is too high, even a good locomotion policy might struggle to maintain stability during dynamic contact events, so I'm curious if they found any specific thresholds for acceptable performance.

Rosa: They did suggest that the continuous nature of the recurrent estimator helps mitigate those latency issues by leveraging proprioception and policy actions to keep tracking going even when vision is limited.

Dev: That’s a crucial point; it means the architecture is designed to be inherently fault-tolerant regarding sensory input, which is something we need to build into our next generation of controllers.

Taro: It really shows that the combination of hierarchical RL with a recurrent belief system is a powerful way to tackle complex, contact-rich interactions where failure is inevitable and recovery is key.

The paper's improvements: Taro: So, to sum up the improvements, AdaptManip isn't just about adding components; it’s about building this structured three-stage strategy—navigation, grasping, and carrying—and then training the whole thing end-to-end using reinforcement learning without any human help.

Rosa: That's right; they are not relying on someone showing the robot exactly how to walk or how to lift a box; the AI learns all those complex behaviors through trial and error guided by a carefully designed reward system.

Dev: What I find particularly interesting is how they separate the learning into two distinct policies: a base locomotion policy for stable walking, and then a residual manipulation policy that only focuses on adapting to the object once it's grasped.

Taro: That hierarchical structure seems smart because it lets the robot focus its learning efforts where they matter most, ensuring it maintains fundamental stability while simultaneously learning the fine motor skills of grasping and delivering.

Rosa: And I think the real muscle here is that recurrent object state estimator, which acts like a persistent memory for the object's location, allowing the system to recover even when visual input gets messy or temporarily missing.

Dev: That memory function is critical because it enables that closed-loop recovery you mentioned; without that continuous update on what the box is doing, any momentary slip would probably cause a complete failure of the manipulation task.

Taro: It means the system can actually handle failures in a dynamic way, instead of just stopping or freezing when something goes wrong during carrying.

Rosa: Exactly, and when we look at sim-to-real transfer, they suggest a specific method where they train the policy using noisy or masked ground truth poses to explicitly model those visual estimation errors and occlusions from the start.

Dev: That explicit modeling of uncertainty in training is smart because it teaches the policy how to behave predictably when its sensor data isn't perfect in reality, which is vital for deployment.

Taro: It suggests a new way of approaching sim-to-real transfer where you don't just hope it works; you train the system to be robust against the specific types of errors it will encounter in the real world.

Rosa: So, this framework moves beyond just achieving a single task; it’s building an adaptable robot capable of performing a sequence of complex, multi-step actions autonomously in an unstructured setting.

Dev: The implication for control systems is that we can design controllers that are inherently resilient to estimation noise by integrating recurrent feedback mechanisms directly into the policy's state space, rather than trying to filter it out externally.

Taro: If this approach scales, I imagine we could see agents capable of performing much more intricate tasks in real-world environments where human supervision isn't feasible at all.

Conclusion: Rosa: So we've seen how AdaptManip uses a combination of recurrent state estimation and hierarchical reinforcement learning to let humanoids autonomously navigate, pick up, and deliver objects without any human guidance or pre-recorded demonstrations.

Dev: That’s right; it really shows how combining robust estimation with adaptive control policies can handle the complexity of whole-body manipulation in real environments.

Taro: I think the biggest impact is how this framework addresses the "what if" scenarios; it’s not just about following a path, but about maintaining competence when unexpected disturbances happen during that delivery.

Rosa: It really is; we're looking at a future where robots aren't just programmed for specific routes but are truly capable of adapting their whole strategy in real-time based on what they perceive.

Dev: From an engineering standpoint, the robustness against failure modes is key here; that online estimation loop means the system can handle transient errors and continue its mission instead of crashing immediately.

Taro: And that ability to recover during manipulation suggests we might see agents capable of handling much more intricate tasks in unstructured settings where human supervision simply isn't available.

Rosa: It’s certainly exciting to think about robots that can operate effectively outside the lab for extended periods, navigating and completing complex tasks on their own using only onboard sensing.

Dev: We've seen strong results on physical hardware, which gives us confidence that the simulation-to-real transfer isn't just a lucky fluke but is based on sound mechanisms.

Taro: That validation across different environments shows that the generalization achieved by this recurrent estimator is quite solid for complex locomotion and manipulation.

Rosa: So, to wrap up, AdaptManip provides a powerful learning-based framework for whole-body humanoid loco-manipulation that tackles navigation, lifting, and delivery autonomously using onboard sensing.

Dev: It sets a high bar for how we design policies that prioritize stability while allowing the system to adapt its behavior based on real-time environmental feedback.

Taro: We've seen how this architecture can handle failures gracefully during complex interaction sequences, which is where the true autonomy lies in field robotics.

cs.RO, cs.LG

Submitted: 2026-02-16

Updated: 2026-09-30

Comments: Website: https://morganbyrd03.github.io/adaptmanip/

Code: https://github.com/duckietown/lib-dt-apriltags

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 91/100

The gist: AdaptManip presents a fully autonomous framework for humanoid robots to perform integrated navigation, object lifting, and delivery by training a robust loco-manipulation policy via reinforcement

Key concepts

Recurrent Object State Estimator
This module is designed to track the manipulated object's 6D pose in real-time, even when the robot has limited vision or experiences occlusions. It fuses data from visual sensors (like cameras), LiDAR, and internal robot signals (proprioception) using a V-LSTM and MLP structure. This allows the system to maintain a continuous belief about where the object is located during manipulation.
Whole-Body Base Policy
This is the foundational reinforcement learning policy trained to generate stable bipedal locomotion for the humanoid robot. It learns how to walk and move its base effectively, providing a stable platform for subsequent manipulation tasks. Its training focuses on tracking base velocity and maintaining joint stability under various physical conditions.
Manipulation Residual Policy
This policy is trained on top of the locomotion policy to learn task-specific actions needed for grasping and lifting objects. It takes the estimated object pose as input, allowing it to adapt its movements while ensuring that the overall robot motion remains stable. This component focuses specifically on how to interact with and move the target object.
Sim-to-Real Transfer
This refers to the ability of a model trained in a simulated environment (like IsaacLab) to perform well when deployed on physical hardware. AdaptManip demonstrates strong sim-to-real transfer by successfully performing tasks in MuJoCo and achieving real-world success, proving the learned policies generalize effectively from simulation to physical robots.

Terminology

Summary

AdaptManip presents a fully autonomous framework for humanoid robots to perform integrated navigation, object lifting, and delivery by training a robust loco-manipulation policy via reinforcement learning without human demonstrations or teleoperation data.

The gist

AdaptManip is a learning-based framework for whole-body humanoid loco-manipulation that autonomously accomplishes navigation, lifting, and delivery through a structured three-stage strategy.

Framework Components and Strategy

The proposed framework consists of three coupled components designed to achieve robust and recovery-capable loco-manipulation:

  1. A recurrent object state estimator that tracks the manipulated object in real time under limited field-of-view and occlusions.

  2. A whole-body base policy for robust locomotion with residual manipulation control for stable object lifting and delivery.

  3. A LiDAR-based robot global position estimator that provides drift-robust localization.

The robot operates through three coordinated stages:

  1. Navigation: The humanoid approaches the target object using LiDAR-based robot pose odometry and proprioceptive feedback, enabling fully onboard localization.

  2. Grasping and Lifting: An online object state estimator guides grasping and coordinated whole-body manipulation, utilizing vision-based sensing to refine the object pose for accurate grasping.

  3. Carrying to Destination: The robot transports the object while maintaining balance under changing contact conditions, with visual information becoming unreliable due to occlusion.

Learning Architecture

The learning process is structured into three training components:

  1. Base Whole-Body Locomotion Policy: This policy is trained using RL to generate stable bipedal locomotion, serving as a fixed foundation for subsequent whole-body manipulation learning. The observation space includes joint positions, velocities, base angular velocity, projected gravity, commanded planar velocity, and base yaw rate.

  2. Manipulation Residual Policy: This policy is trained on top of the base policy to learn task-specific adaptations for object grasping and lifting while preserving locomotion stability. Its observation space is augmented with the estimated 6D box pose and previous action at−1.

  3. Recurrent Online Object State Estimator: This module fuses visual observations (using a V-LSTM and MLP) and proprioceptive signals to infer the object pose online, aiming to achieve robust manipulation under partial or missing visual inputs.

Reward Design and Training Details

The reward design is weighted to encourage locomotion stability, gait shaping, motion regularization, and constraint violation penalties. The locomotion reward includes terms for command tracking (e.g., tracking base linear velocity), gait shaping (penalizing joint accelerations), motion regularization (penalizing joint velocities and torques), and constraint violation penalties related to joint limits. For the manipulation policy, the reward is augmented with manipulation-specific objectives, including kinematic tracking, box stabilization, contact force quality, and slip avoidance terms such as contact-related terms penalize excessive relative motion between the robot and the box and discouraging tangential hand motion indicative of slipping.

Evaluation and Results

Experimental results show that AdaptManip significantly outperforms baseline methods in adaptability and overall success rate. In simulation experiments using IsaacLab, AdaptManip achieved an 85% success rate, comparable to Pure RL + FK (88%) and Oracle (91%), indicating near-optimal performance when reliable object state information is available. Crucially, during sim-to-sim transfer evaluation in the unseen MuJoCo environment, AdaptManip achieved a 75% success rate, which was comparable to the 79% of Oracle. The paper further demonstrates that the recurrent estimator's advantage is evident in hardware experiments: while visual estimates degrade during floating-base motion (e.g., walking), the recurrent estimator continues to track the object by leveraging robot proprioception and policy actions, confirming effective sim-to-real transfer in a zero-shot manner. The system successfully demonstrated fully autonomous real-world navigation, object lifting, and delivery on a Unitree G1 humanoid robot using only onboard sensing.

Key Contributions

The main contributions are:

  1. Introducing AdaptManip, a learning-based framework for whole-body humanoid loco-manipulation that autonomously accomplishes navigation, lifting, and delivery through a structured three-stage strategy.

  2. Developing an online, recurrent object state estimation module that fuses LiDAR, vision, and proprioceptive sensing to enable robust and recovery-capable loco-manipulation using only onboard sensors.

  3. Validating the effectiveness of AdaptManip through extensive simulation studies and real-world experiments on physical humanoid hardware, showing effective sim-to-real transfer.

Conclusion

AdaptManip successfully combines multi-modal inputs of LiDAR, vision, and proprioception to maintain a recurrent belief of the box pose and hierarchical RL in order to effectively learn a policy which utilizes the pose for picking up and carrying a box from an initial position to a target location. Future work could include trying more extended tasks or incorporating additional sensor modalities.


How it works

The framework is structured into three coupled components:

Improvements for AI systems

Here are the specific improvements to AI systems inspired by AdaptManip, and what those improved systems can achieve:

  1. A fully autonomous, zero-shot whole-body humanoid robot capable of performing complex pick-and-place tasks (navigation, lifting, and delivery) in unstructured real-world environments without any prior human demonstrations or teleoperation data.

  2. An object state estimation module that fuses multimodal inputs (LiDAR odometry, vision from a camera/AprilTag, and proprioceptive feedback) to provide robust 6D pose tracking of an object even when it is partially occluded by the robot's body or hands during manipulation.

  3. A hierarchical reinforcement learning framework where a base locomotion policy ensures stable bipedal walking, while a residual manipulation policy adaptively learns contact-rich whole-body actions (grasping and lifting) by leveraging the real-time estimated object state, ensuring both mobility and task success under failure conditions (e.g., slippage).

  4. A closed-loop control system that actively recovers from transient failures during manipulation, such as losing grasp stability, by executing corrective arm motions guided by the recurrent state estimator's continuous belief update on the object's pose.

  5. An improved sim-to-real transfer methodology where the policy is trained in simulation using noisy/masked ground-truth poses to explicitly model visual estimation errors and occlusions, leading to higher success rates (e.g., 75% success rate) when deployed on real hardware without further fine-tuning.

Sources

Related papers