Reinforcement Learning for Real-Time Vision-Language-Action Policies
cs.RO, cs.LG
Submitted: 2026-09-16
Updated: 2026-09-16
Project page: https://pd-perry.github.io/real-time-expo-ft
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reinforcement learning fine-tuning on top of large, pretrained Vision-Language-Action (VLA) models offers promise for highly reliable robot deployment.
Terminology
Abstract
Reinforcement learning fine-tuning on top of large, pretrained Vision-Language-Action (VLA) models offers promise for highly reliable robot deployment. However, because of their scale, modern VLA models suffer from high inference latency, so the observation used to select an action is often stale by execution time, creating a distribution shift that can substantially degrade reliability and performance. Prior work has explored asynchronous policy execution to reduce the effect of latency, but these methods are mostly built on imitation learning and offer no mechanism for moving beyond the training distribution toward higher reliability. We close this gap by enabling RL fine-tuning that meets the real-time control requirements of dynamic real-world manipulation. Our approach builds on EXPO-FT, a framework for sample-efficient, reliable VLA fine-tuning with reinforcement learning, and decouples slow, expressive action generation from fast, reactive action edits: a large pretrained VLA proposes action chunks using its strong behavior prior, while a lightweight edit policy performs fast, reactive decision-making by editing actions in response to changes in state, conditioned on the latest observation. We instantiate this as Real-Time EXPO-FT, an RL framework for finetuning real-time VLA policies. On the Kinetix benchmark, Real-Time EXPO-FT enables a delayed policy to achieve the best performance among delayed and non-delayed methods in 10 out of 10 environments. On four dynamic real-world tasks, robot object passing, ball balancing, table soccer kicking, and dynamic object picking, with online robot data capped at 10 minutes, Real-Time EXPO-FT improves average policy performance from 42% to 97%, all without human intervention, demonstrating rapid, sample-efficient adaptation to challenging real-world dynamics. Website: https://pd-perry.github.io/real-time-expo-ft
Sources
- End-to-End Training of Deep Visuomotor Policies
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Dexterous Manipulation with Deep Reinforcement Learning: Efficient, General, and Low-Cost
- Soft Actor-Critic Algorithms and Applications
- IRIS: Implicit Reinforcement without Interaction at Scale for Learning Control from Offline Robot Manipulation Data
- Self-Improving Robots: End-to-End Autonomous Visuomotor Reinforcement Learning
- Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
- Residual Off-Policy RL for Finetuning Behavior Cloning Policies
- ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
- Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning
- Running VLAs at Real-time Speed
- Leave No Observation Behind: Real-time Correction for VLA Action Chunks
- Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- $\pi_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models
- EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models
- TQL: Scaling Q-Functions with Transformers by Preventing Attention Collapse
- Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving