Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

arXiv:2608.12063 · cs.RO, cs.AI · Submitted 2026-08-12 · Read on arXiv

Martin Schuck, Maks Sorokin, Simone Manni, Duy Ta, Angela P. Schoellig, Marco Hutter, Simon Le Cleac'H, Jan Brüdigam

RAI Institute · Technical University of Munich · ETH Zurich

cs.RO, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Project page: https://pages.rai-inst.com/smpc2rl

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: This paper presents a framework for learning loco-manipulation policies using Sample-based Model Predictive Control (SMPC) demonstrations with sparse offline-to-online Reinforcement Learning.

Terminology

Summary

This paper presents a framework for learning loco-manipulation policies using Sample-based Model Predictive Control (SMPC) demonstrations with sparse offline-to-online Reinforcement Learning. The authors propose leveraging SMPC entirely in simulation as an automated, rapidly tunable expert to generate massive offline datasets, which solves the fundamental exploration problem, allowing training of an off-policy RL agent using purely sparse task rewards.

The paper's contributions are threefold: (1) a tuning-free RL training framework that leverages SMPC as an easily tunable automated data generator in simulation, enabling sparse-reward off-policy RL and circumventing manual reward-tuning bottlenecks; (2) demonstration that bootstrapping sparse off-policy RL with SMPC datasets yields hardware-deployable policies that surpass the original optimal control teacher in performance; (3) validation across vastly different morphologies, deploying learned skills on real hardware for both an arm-equipped Spot quadruped and a G1 humanoid.

The approach uses a hierarchical control architecture that decouples task-level navigation and manipulation from base maneuvering. The high-level policy's action space is defined as a high = [Δv cmd, Δq cmd arm, Δh cmd, Δp cmd], representing deltas to current desired planar base velocity, arm joint positions, torso height target, and torso pitch target. The low-level stabilization controller is a whole-body maneuvering controller trained via the ReLIC framework, which is frozen for all downstream task learning.

The training uses a strictly sparse task reward function: r = 0 at goal, r = -2/(1-γ) if robot crashed, and r = -1 otherwise, where γ is the discount factor of 0.99. The crash penalty is chosen such that agents prefer constant negative reward over crashing.

The authors employ an offline-to-online training strategy using a modified FastTD3 architecture. During initial learning phases, 50% of transitions in the replay buffer are replaced with pre-collected expert data. A curriculum phases out expert data once the agent achieves a sufficient empirical success rate, shifting to pure online learning.

The SMPC data collection uses vectorized MuJoCo Warp environments with a massively parallelized tiled approach. The SMPC achieves approximately 0.5× real-time performance on a single RTX 5090, allowing reward tuning in minutes. The architecture generates high-quality demonstration datasets at a rate of one million samples per hour, providing offline data required to bootstrap sparse-reward RL agents within four hours on a single GPU.

The framework was validated on five diverse tasks: Spot reach, Spot box push, Spot tire upright, Spot tire roll, and G1 box push. Across all tasks, policies achieved near-perfect success rates with little variance between different seeds. The policies trained exclusively on sparse rewards were reliably deployed on physical hardware across all five tasks and both robotic platforms.

The paper answers four research questions through structured ablations:

Q1 (Improving over seen expert data): Sparse-reward policies consistently outperform SMPC experts, with some tasks showing improvements above 50%. Learned policies also show increased consistency with 11-45% reduction in standard deviation of task duration.

Q2 (Amount of data): Required dataset size correlates with task complexity. While easier tasks can be solved with little data, more complex whole-body manipulation tasks require significantly larger datasets. The most difficult task requires four million samples collected within 4 GPU hours.

Q3 (Quality of data): Most tasks are relatively robust to lower-quality data, but tasks requiring high coordination (like tire rolling) are sensitive to data quality.

Q4 (Multimodality): Multi-modal SMPC data heavily affects training, and agents fail to learn viable policies. Enforcing a single behavioral mode in demonstration data is required for successful offline initialization.

The authors acknowledge that achieved behavioral optimality remains local, as the policy is tied to the dataset's distribution. The current control architecture is limited by frozen weights of the low-level controller, preventing adaptation to task-specific physical disturbances. Real-world deployment relies on state-based information, requiring training on vision-based data or distillation into vision-based policies for unstructured environments.

Improvements for AI systems

Improvement 1: Automated Expert-Guided Exploration for Sparse-Reward RL

I can integrate a sample-based MPC (SMPC) as a tunable, simulation-only data generator to bootstrap off-policy RL agents. This eliminates the need for manual reward shaping or dense reward engineering. The improved AI system can autonomously collect millions of high-quality demonstrations in hours (e.g., 1M samples/hour on a single GPU) and then train a sparse-reward policy that outperforms the expert by 50%+ in task success and reduces task duration variance by 11–45%. It can be deployed directly on diverse robot morphologies (quadrupeds with arms, humanoids) without per-task reward tuning.

Improvement 2: Hierarchical Control with Frozen Low-Level Stabilization

I can adopt a two-level architecture: a high-level policy outputs deltas for base velocity, arm joint positions, torso height, and pitch, while a pre-trained whole-body controller (e.g., ReLIC) handles low-level stabilization. The improved AI system can learn new manipulation tasks (reach, push, tire upright/roll, box push) with near-perfect success rates across five tasks and two hardware platforms, while keeping the low-level controller frozen—reducing training complexity and enabling rapid task transfer.

Improvement 3: Curriculum-Based Offline-to-Online Data Mixing

I can implement a replay buffer strategy where 50% of transitions are expert SMPC data during early training, then phase out expert data once empirical success rate rises. The improved AI system can transition smoothly from offline pretraining to pure online RL, avoiding catastrophic forgetting and achieving hardware-deployable policies within 4 GPU hours, even for complex whole-body tasks requiring 4M samples.

Improvement 4: Data Quality and Mode Enforcement for Robust Learning

I can filter SMPC demonstrations to enforce a single behavioral mode (e.g., consistent tire-rolling direction) and ensure data quality thresholds for coordination-heavy tasks. The improved AI system can avoid failure modes from multimodal data (which otherwise prevents viable policy learning) and remain robust to lower-quality data for simpler tasks, while maintaining high performance on tasks requiring precise coordination.

Improved AI System Capabilities:

  • Learns complex loco-manipulation skills from sparse rewards only, using automated SMPC-generated datasets in simulation.

  • Deploys policies on real robots (Spot + arm, G1 humanoid) for tasks like box pushing and tire uprighting with near-perfect success and low variance.

  • Adapts to new tasks in hours without manual reward tuning, using a single GPU and vectorized MuJoCo environments.

  • Provides a scalable pipeline for other embodied AI tasks requiring whole-body coordination, with explicit control over data quantity, quality, and multimodality.

Sources

Related papers