ReCo: Response-Consistent Locomotion with Policy-Aware MPC for Legged Manipulation

arXiv:2610.01612 · cs.RO · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "ReCo: Response-Consistent Locomotion with Policy-Aware MPC for Legged Manipulation".

Rosa: Continuous legged manipulation requires accurate end-effector tracking while the base keeps walking, and this paper presents ReCo, a framework that couples response-consistent locomotion with policy-aware MPC for legged manipulation.

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: We’ve been discussing the mechanics of ReCo, but I want to start by talking about the paper's title and who came up with it; it’s "ReCo: Response-Consistent Locomotion with Policy-Aware MPC for Legged Manipulation." It tells you exactly what this system is designed to do: make locomotion consistent and use policy-aware MPC for arm coordination.

Dev: I agree, the name itself suggests a focus on consistency, which is crucial when you're trying to coordinate two very different motions like walking and reaching. The authors are Kuankuan Sima, Yichao Gao, Chenxi Gu, Kefan Zhao, and Lin Zhao from National University of Singapore who developed this framework.

Taro: I was looking at their background in autonomous systems research; they seem to be drawing on a mix of reinforcement learning techniques and model predictive control approaches to solve these kind of problems. It sounds like a solid combination for achieving complex locomotion goals.

Rosa: That combination is what caught my attention; combining RL for the policy's decision-making with MPC for the low-level coordination is a proven path, but ReCo seems to refine how they link those two parts together specifically for legged manipulation.

Dev: They are essentially tackling the difficulty that standard MPC struggles with because it can only predict base motion that it can foresee, whereas a learned policy's behavior changes based on things like gait phase or contact events.

Taro: That’s the core tension they’re addressing; the policy is dynamic and context-dependent, but the planner needs a predictable model to work with, and ReCo seems to bridge that gap by explicitly modeling that response.

Rosa: So, when we look at their overall ambition here, it seems they are trying to create a system where the base locomotion doesn't just happen, but happens in a way that is repeatable and controllable across different scenarios.

Dev: I think the implication for control engineering is that if you can make the policy's command response repeatable through reference tracking, you gain a much more stable foundation for planning commands. It reduces the need for overly aggressive or reactive safety margins in the MPC formulation itself.

Taro: And from an autonomy perspective, it means we are building policies that are not just capable of performing a task once in simulation, but ones that have a repeatable execution profile when deployed in uncertain physical environments.

Rosa: That’s a big step toward making legged systems truly reliable for long-duration missions where failure due to erratic locomotion is something we have to avoid. It sounds like they are aiming for robustness through training structure rather than just brute-force tuning during deployment.

Dev: So, the title really sets the stage: it's not just about walking or reaching; it’s about making those two intertwined movements happen in a predictable, coordinated way using this novel response shaping and MPC coupling.

The paper's summary: Rosa: Now we’re getting into the substance of what ReCo actually proposes; the authors summarize it as a framework that couples response-consistent locomotion with policy-aware MPC to solve continuous legged manipulation problems. Essentially, they are using response shaping to train the locomotion policy to be consistent and repeatable across randomized dynamics.

Dev: That training method involves driving five command channels—planar velocity, yaw rate, height, pitch, and roll—with critically damped reference generators that enforce target responses against the actual robot movement. It sounds like a very explicit way of telling the AI how it *should* react under different conditions.

Taro: The model identification part is also key here; they identify a closed-loop response model that captures how the policy and robot interact under candidate commands, including gait-periodic base motion modeled as a harmonic series.

Rosa: That specific harmonic series model for the base height, delta j = N h / sum n=one (a zero j n + a one j n nu) (n phi + beta jn), is quite detailed, suggesting a deep dive into modeling the periodic nature of the gait itself.

Dev: That level of detail in modeling the base dynamics suggests they've done a lot of work to capture the empirical closed-loop response without necessarily needing to model every single arm-induced wrench explicitly during that identification phase.

Taro: So, in summary, ReCo uses response shaping for training consistency and then feeds this identified model into a policy-aware MPC that jointly plans locomotion commands and arm motion. It’s a very integrated system architecture.

Rosa: It really seems like they’ve managed to create a tight feedback loop where the learned policy informs the planner, and the planner uses a model of that interaction to make sure everything stays coordinated during manipulation.

Dev: I think what's important is how this coupling allows the arm to anticipate things like command lag and gait oscillations, which are inherently dynamic issues that simple models often fail to capture on their own.

The paper's improvements: Rosa: Focusing on what they actually improved, the paper highlights several key areas where ReCo makes a difference; they point out the training method involving proximal policy optimization and a gait-conditioned interface like Walk These Ways.

Dev: They also emphasize adding response shaping terms to augment the locomotion reward function, combining positive terms r+ zero and signed penalties r- zero into a combined reward structure r = r+ (c - r-) where c > zero. This is a powerful way to encourage that repeatability.

Taro: The cross-domain consistency enforcement, penalizing deviations between randomized instances and the nominal response using an instance d incurs loss term, also seems like a necessary step to ensure the trained policy generalizes well beyond the specific set of dynamics it was trained on.

Rosa: And then there's the identified closed-loop response model itself; this model acts as a predictive interface for the MPC, allowing it to jointly plan locomotion and arm motion based on that specific dynamic knowledge.

Dev: The MPC stage cost they use is quite comprehensive, including terms for position error e p squared Q p, orientation error e R squared Q R, and command–response discrepancies like W E - ref WE squared Q v.

Taro: Those specific terms in the cost function show they are explicitly penalizing those command-response discrepancies, which directly ties back to their response shaping goals and makes the MPC directly responsible for managing those issues.

Rosa: The results speak for themselves; they report that ReCo reduces position and orientation root-meansquare error (RMSE) by twenty-eight point seven percent and twenty-seven point four percent relative to the best baseline for each metric, which is a significant quantitative improvement in tracking accuracy.

Dev: On top of that, response shaping lowered the normalized prediction mean squared error by fifty-nine point six percent, which shows how much better the model is at predicting future states under these conditions compared to other methods.

Conclusion: Rosa: So, to wrap up our discussion on ReCo: Response-Consistent Locomotion with Policy-Aware MPC for Legged Manipulation, it seems the paper presents a method that systematically trains locomotion policies for consistency and then uses an identified response model within a policy-aware MPC for joint planning.

Dev: Exactly; the training involves careful use of response shaping and reward augmentation, followed by fitting a closed-loop response model to inform the MPC's predictions about command lag and gait oscillations. It’s a complete control loop where locomotion and manipulation are tightly coupled through this predictive interface.

Taro: The main implication I see is that we're developing methodologies for injecting predictability into learned locomotion policies, which could be useful for any autonomous system relying on reinforcement learning to perform complex physical actions in the real world.

Rosa: It really seems like this work lays a solid foundation by providing concrete improvements in tracking error and showing that coordinated base and arm motion can actually be achieved onboard. We'll keep an eye on how they address those limitations we discussed, especially regarding online adaptation as we move toward more complex scenarios.

Dev: I think the next big hurdle for this architecture is ensuring that the identification of the response model remains accurate under significant external disturbances or when the robot encounters truly unexpected dynamics outside its training distribution.

Taro: I'd say that while it achieves excellent performance in simulation and hardware tests, we still need to see if this level of coordination holds up when faced with unpredictable, unstructured interactions from a completely novel environment.

Rosa: Well, before we sign off on ReCo: Response-Consistent Locomotion with Policy-Aware MPC for Legged Manipulation, it’s clear they've provided a very structured approach to making legged manipulation more predictable and coordinated.

Dev: Indeed; the combination of response shaping and policy-aware MPC gives us a powerful tool for managing the inherent complexities of moving manipulators.

Taro: I think this framework moves us closer to deploying these systems in situations where robust, continuous physical interaction is required, which is a big step for autonomy.

Kuankuan Sima, Yichao Gao, Chenxi Gu, Kefan Zhao, Lin Zhao

National University of Singapore

cs.RO

Submitted: 2026-10-01

Updated: 2026-10-01

Code: https://github.com/leggedrobotics/ocs2

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 89/100

The gist: Continuous legged manipulation requires accurate end-effector tracking while the base keeps walking, and this paper presents ReCo, a framework that couples response-consistent locomotion with

Key concepts

Response Shaping
This training method modifies the locomotion policy's reward function by adding positive and negative terms. It encourages the policy to produce consistent responses across different random dynamics by enforcing target responses against actual outputs using critically damped reference generators.
Closed-Loop Response Model
This model captures how the robot's dynamics respond to commands, incorporating both command-response dynamics and gait-periodic base motion. It is identified offline using least squares on prediction errors to create a predictive interface that accounts for the system's actual closed-loop behavior.
Policy-Aware MPC
This component plans both the robot's walking commands and the arm's motion simultaneously by using the identified response model and arm kinematics. The MPC minimizes a cost function that includes position error, orientation error, and discrepancies between commanded responses to anticipate lag.

Terminology

Summary

Continuous legged manipulation requires accurate end-effector tracking while the base keeps walking, and this paper presents ReCo, a framework that couples response-consistent locomotion with policy-aware MPC for legged manipulation. The core contribution is a training method that makes the command response of a locomotion policy repeatable through reference track ing, allowing an identified closed-loop response model to be used by MPC to jointly plan locomotion commands and arm motion.

How it works

The framework couples response-consistent locomotion with policy-aware MPC using three main components: Response shaping, a closed-loop response model, and a policy-aware MPC. Response shaping trains the locomotion policy to respond consistently and repeatably across randomized dynamics by encouraging repeatable command transients and gait-phase-dependent body motion. This is achieved by driving five command channels—planar velocity, yaw rate, height, pitch, and roll—with critically damped reference generators that enforce target responses against the actual response.

How it works

The closed-loop response model captures the dynamics of the policy–robot closed loop under candidate commands. It combines command-response dynamics with gait-periodic base motion. Specifically, for base height (z), a speed-dependent gait-periodic component is modeled as a harmonic series: a harmonic series: δj = Nh / ∑ n=1 (a0 jn +a1 jnν)sin(nφ +βjn). This model is identified offline using bounded nonlinear least squares on multi-step open-loop prediction errors, capturing the empirical closed-loop response without explicitly modeling arm-induced wrenches.

How it works

The policy-aware MPC jointly plans locomotion commands and arm motion by utilizing the identified response model and arm kinematics. The MPC solves a constrained optimization problem to minimize a stage cost that includes position error, orientation error, and command–response discrepancies: The stage cost is l = 1/2 ep squared Qp + 1/2 eR squared QR + 1/2 p˙WE − p˙ref WE squared Qv + 1/2 u squared Ru +lbase +lposture. This coupling allows the arm to anticipate command lag and gait oscillations.

How it works

The training process involves several steps to ensure robustness. First, the locomotion policy is trained using proximal policy optimization and a gait-conditioned interface. Second, response shaping terms are added to augment the locomotion reward: "Positive terms r+ ≥ 0 and signed penalties r− ≤ 0 are combined as r = r+ exp(c−r−), with c− > 0 [4]. Third, cross-domain consistency is enforced by penalizing deviations between randomized instances and the nominal response: Instance d incurs l d domain = m d / m d y˜ d −y˜ nom squared W d". Finally, model identification is performed offline to fit the closed-loop response model to each frozen policy.

How it works

The experimental evaluation compares ReCo against several baselines, including RL+MPC (Ma et al.), Pure MPC, and learned whole-body control (RoboDuet). On the simulation benchmark, ReCo reduces position and orientation root-meansquare error (RMSE) by 28.7% and 27.4% relative to the best baseline for each metric. Furthermore, response shaping lowers normalized prediction mean squared error (MSE) by 59.6%, and hardware experiments demonstrate onboard execution of continuous legged manipulation with coordinated base and arm motion. The framework successfully achieves tracking improvements while maintaining high success rates in challenging conditions.

The gist

ReCo reduces position and orientation root-meansquare error (RMSE) by 28.7% and 27.4% relative to the best baseline for each metric, response shaping lowers normalized prediction mean squared error (MSE) by 59.6%, and hardware experiments demonstrate onboard continuous legged manipulation with coordinated base and arm motion.

Key Contributions Enumerated:

  1. A training method that makes the command response of a locomotion policy repeatable through reference track ing, ing, phase repeatability, steady-state gain, and crossdomain consistency rewards.

  2. A policy-aware MPC that jointly plans locomotion commands and arm motion for continuous world-frame EE SE(3) tracking, where an identified closed-loop response model with a gait-periodic component serves as the predictive interface to the learned policy.

  3. Simulation and hardware experiments in which ReCo reduces position and orientation root-meansquare error (RMSE) by 28.7% and 27.4% relative to the best of five baselines for each metric, response shaping lowers normalized prediction mean squared error (MSE) by 59.6%, and the full system runs onboard.

  4. Response shaping lowers position RMSE from 0.

Improvements for AI systems

Here are the specific improvements that can be made to existing AI systems, based on the principles and results presented in ReCo: Response-Consistent Locomotion with Policy-Aware MPC for Legged Manipulation, and what these improved systems can achieve.


  1. The system should incorporate a training methodology called Response Shaping into its locomotion policy.

  2. This response shaping involves training the policy to produce consistent command responses across randomized dynamics, specifically by using regularization terms that penalize deviations in:

  3. Planar velocity, yaw rate, height (relative to nominal), and pitch commands during gait phases.

  4. The system should also incorporate a mechanism for Model Identification of the closed-loop response model of the locomotion policy. This involves fitting a low-order model to the policy's command response dynamics, explicitly modeling both commanded inputs and resulting base motion (including gait-periodic height and attitude variations).

  5. The system should utilize a Policy-Aware Model Predictive Control (MPC) architecture that jointly plans locomotion commands and arm motion.

  6. The MPC should use the identified closed-loop response model as its predictive interface to the learned policy, allowing it to anticipate command lag and gait oscillations that simple velocity models miss.

  7. The system should be capable of performing continuous world-frame end-effector (EE) tracking while simultaneously maintaining robust locomotion, even when the base motion is non-deterministic or subject to contact uncertainty.

This improved AI system can perform the following specific tasks:

  1. It can execute complex, sustained manipulation tasks (e.g., inspection, scanning, tool positioning) while walking across uneven or uncertain terrain without significant drift in object tracking.

  2. It will demonstrate superior coordination between base locomotion and arm motion, leading to significantly higher success rates (up to 75/100 runs compared to baselines) and lower tracking errors (position RMSE reduced by 28.7% and orientation RMSE reduced by 27.4%).

  3. It can maintain high performance under external disturbances (e.g., applied forces or pushes), exhibiting minimal increase in tracking error, unlike baseline methods where error often spikes significantly under impulse application.

  4. It can operate onboard physical hardware, achieving continuous manipulation tasks in real-world environments with verified accuracy (position RMSE of 0.153 m and orientation RMSE of 9.93°) when paired with appropriate state estimation and feedback mechanisms like motion capture or odometry fusion.

Abstract

Continuous legged manipulation requires accurate end-effector tracking while the base keeps walking. Combining reinforcement learning (RL) with model predictive control (MPC) suits this task: the learned policy provides robust locomotion, while MPC coordinates the base and arm to compensate for tracking errors. However, MPC can compensate only for base motion that it can predict, and a learned policy's command response varies with gait phase, contact, and payload. We present ReCo, a framework that couples response-consistent locomotion with policy-aware MPC for legged manipulation. Response shaping trains the policy to respond to commands consistently and repeatably across randomized dynamics. An identified closed-loop response model then lets MPC jointly plan locomotion commands and arm motion. On the simulation benchmark, ReCo reduces position and orientation root-mean-square error (RMSE) by 28.7% and 27.4% relative to the best baseline for each metric. Real-world experiments demonstrate onboard continuous legged manipulation with coordinated base and arm motion.

Sources

Related papers