Prompt-Driven Exploration: Language as an Exploration Space for VLA Reinforcement Learning
cs.LG, cs.AI
Submitted: 2026-07-09
Updated: 2026-09-18
Code: https://github.com/haosulab/mplib
Project page: https://xinyunsunshine.github.io/prompt-rl
License: http://creativecommons.org/licenses/by/4.0/
The gist: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers.
Terminology
Abstract
Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject stochasticity in the action space, but such jitter only yields rollouts close to the original. Escaping a weak policy often requires global perturbations that action noise cannot produce. Large language models (LLMs) and vision-language-action (VLA) models offer a pathway: they condition the policy on a natural language prompt, and since the rollout follows from it, modifying the prompt induces global changes. The challenge is finding prompts that induce useful global changes. With a weak policy that rarely succeeds, reward is too sparse to select on. Our idea is to refine prompts from the rollouts themselves: a vision-language model (VLM) reasons over the rollout video, diagnoses how the policy responded, and rewrites the prompt to elicit better behavior next time. This procedure resembles posterior sampling, a classical RL exploration framework, at the level of prompts: the VLM maintains an implicit distribution over useful prompts and updates it from observed rollouts. We call this strategy Prompt-Driven Exploration (PDE). Across manipulation and reasoning tasks, PDE enables RL to learn successful policies even from zero-reward starts, and improves sample efficiency more broadly. Our website is available at https://xinyunsunshine.github.io/prompt-rl.
Sources
- GPT-4 Technical Report
- What learning algorithm is in-context learning? Investigations with linear models
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- Exploration by Random Network Distillation
- $\pi_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models
- Curiosity-driven Red-teaming for Large Language Models
- GPT-4o System Card
- OpenAI o1 System Card
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Embodied Red Teaming for Auditing Robotic Foundation Models
- OpenVLA: An Open-Source Vision-Language-Action Model
- Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
- LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
- VLGOR: Visual-Language Knowledge Guided Offline Reinforcement Learning for Generalizable Agents
- VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Interactive Post-Training for Vision-Language-Action Models
- ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks