Accelerating Visual Policy Learning with Sampling-Based Model Predictive Control
cs.RO, cs.AI, cs.LG
Submitted: 2026-09-17
Updated: 2026-09-17
Comments: 8 pages, 6 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Learning visual policies for locomotion and manipulation requires coordinating contact with the environment and can incur substantial computation and GPU memory costs.
Terminology
Abstract
Learning visual policies for locomotion and manipulation requires coordinating contact with the environment and can incur substantial computation and GPU memory costs. First-order policy gradients (FoPG) reduce training cost through differentiable simulation, but local optimization can converge to unintended contact patterns. To address this shortfall, we propose Sampling-Guided Policy Search (SGPS), which couples recurring action-target refinement by sampling-based model-predictive control with first-order policy optimization. Behavior cloning initializes the policy from sampled actions; training then alternates sampling-based refinement with short-horizon FoPG updates under perturbed initial states and randomized dynamics. For visual policy training, we use a decoupled FoPG formulation that excludes rendering from the computation graph, enabling direct learning from depth observations without a state-policy teacher. On a single GPU, SGPS learns policies for locomotion, obstacle traversal, crate pushing, and bimanual carrying on simulated Unitree Go2 and G1 robots. Our experiments further show that refinement improves policy learning beyond initialization and tracking alone. For hardware deployment, the distilled policy transfers zero-shot to a real Go2 and uses onboard depth to autonomously trot, crawl, clear hurdles, and switch between these behaviors.
Sources
- Efficient On-policy Visual-RL via Stochastic Decoupled Policy Gradient
- Robot Parkour Learning
- Extreme Parkour with Legged Robots
- ANYmal Parkour: Learning Agile Navigation for Quadrupedal Robots
- Asymptotically Optimal Ergodic Coverage on Generalized Motion Fields
- Proximal Policy Optimization Algorithms
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
- Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning
- Accelerating Visual-Policy Learning through Parallel Differentiable Simulation
- RMA: Rapid Motor Adaptation for Legged Robots
- Learning by Cheating
- Accelerated Policy Learning with Parallel Differentiable Simulation
- Learning Quadruped Locomotion Using Differentiable Simulation
- Stabilizing Reinforcement Learning in Differentiable Multiphysics Simulation
- Predictive Sampling: Real-time Behaviour Synthesis with MuJoCo
- Sample-efficient Cross-Entropy Method for Real-time Planning
- Full-Order Sampling-Based MPC for Torque-Level Locomotion Control via Diffusion-Style Annealing
- Sim-to-Real: Learning Agile Locomotion For Quadruped Robots
- Legged Locomotion in Challenging Terrains using Egocentric Vision
- Walk These Ways: Tuning Robot Control for Generalization with Multiplicity of Behavior
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving