From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention
cs.RO, cs.LG
Submitted: 2026-09-18
Updated: 2026-09-30
Comments: Project page: https://destiny000621.github.io/PARTS/
Project page: https://destiny000621.github.io/PARTS
License: http://creativecommons.org/licenses/by/4.0/
The gist: A pretrained robot foundation policy may execute most of a long-horizon task yet repeatedly fail at a few critical subtasks.
Terminology
Abstract
A pretrained robot foundation policy may execute most of a long-horizon task yet repeatedly fail at a few critical subtasks. Collecting additional full-task demonstrations for supervised fine-tuning (SFT) requires operators to repeat behaviors the policy already performs well. Reinforcement learning (RL) fine-tuning offers a promising path to bridge this gap, but existing approaches struggle to solve long-horizon tasks using only sparse rewards. We present PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world subtask RL framework that concentrates practice at these bottlenecks while allowing training rollouts to proceed with minimal human intervention. The frozen pretrained policy supplies nominal actions throughout execution, while agent-generated selectors and success verifiers activate residual corrections and provide local outcome rewards. These rewards support learning from successful subtasks even when complete-task successes are scarce. Training combines online RL with success-reweighted retraining, and each retrained residual policy is redeployed to collect further experience. Humans identify bottlenecks during setup and perform physical resets when needed. On bimanual YAM and single-arm Franka tasks, PARTS improves complete-task success from 32% to 61% and from 50% to 95%, respectively, using tens of minutes of real-world RL rollouts per task on average. Compared with existing real-world RL fine-tuning methods, PARTS raises full-task success by more than 25% under the same robot-rollout budget while requiring less human involvement.
Sources
- ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
- World Action Models are Zero-shot Policies
- Residual Off-Policy RL for Finetuning Behavior Cloning Policies
- Planning to Practice: Efficient Online Fine-Tuning by Composing Goals in Latent Space
- RL Token: Bootstrapping Online RL with Vision-Language-Action Models
- TORL-VLA: Tactile Guided Online Reinforcement Learning for Contact-Rich Manipulation
- Steering Your Diffusion Policy with Latent Space Reinforcement Learning
- EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models
- Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models
- RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models
- Skill-based Model-based Reinforcement Learning
- OpenVLA: An Open-Source Vision-Language-Action Model
- MoRI: Mixture of RL and IL Experts for Long-Horizon Manipulation Tasks
- BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models
- Data and Learning Where it Matters for Contact-Rich Manipulation
- UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning
- Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids
- RHO: Your Coding Agent is Secretly a Roboticist
- CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
- Offline Retraining for Online RL: Decoupled Policy Learning to Mitigate Exploration Bias
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving