Scaling Sim-to-Real VLA Reinforcement Learning with Generative 3D Worlds
cs.RO, cs.AI, cs.LG
Submitted: 2026-03-19
Updated: 2026-09-29
Comments: Accepted to CoRL 2026. Project page: https://horizonrobotics.github.io/gail/projects/scaling-sim-to-real-rl-vla/
Code: https://github.com/ZiYang-xie/WorldGen
Project page: https://horizonrobotics.github.io/gail/projects/scaling-sim-to-real-rl-vla
License: http://creativecommons.org/licenses/by/4.0/
The gist: The strong performance of large vision-language models (VLMs) trained with reinforcement learning (RL) has motivated similar approaches for fine-tuning vision-language-action (VLA) models in robotics.
Terminology
Abstract
The strong performance of large vision-language models (VLMs) trained with reinforcement learning (RL) has motivated similar approaches for fine-tuning vision-language-action (VLA) models in robotics. Many recent works fine-tune VLAs directly in the real world to avoid addressing the sim-to-real gap. While real-world RL circumvents sim-to-real issues, it inherently limits the generality of the resulting VLA, as scaling scene and object diversity in the physical world is prohibitively difficult. This leads to the paradoxical outcome of transforming a broadly pretrained model into an overfitted, scene-specific policy. Training in simulation can instead provide access to diverse scenes, but designing those scenes is also costly. In this work, we show that VLAs can be RL fine-tuned across broad scene and object distributions and with reduced labor by leveraging 3D world generative models. Using these models together with a language-driven scene designer, we generate 100 diverse interactive scenes containing unique objects and backgrounds, enabling scalable and highly parallel policy learning. Starting from a pretrained imitation baseline, our approach increases simulation success from 9.7% up to 79.8% while achieving a 1.25 times speedup in task completion time. We further demonstrate successful sim-to-real transfer enabled by the quality of the generated scenes together with domain randomization, improving real-world success from 21.7% to 75% and achieving a 1.13 times speedup. Finally, we further highlight the benefits of leveraging the effectively unlimited data from 3D world generative models through an ablation study showing that increasing scene diversity directly improves zero-shot generalization.
Sources
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models
- GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
- $\pi^{*}_{0.6}$: a VLA That Learns From Experience
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
- Proximal Policy Optimization Algorithms
- Diffusion Policy Policy Optimization
- Much Ado About Noising: Dispelling the Myths of Generative Robotic Control
- Boosting Continuous Control with Consistency Policy
- Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- OpenVLA: An Open-Source Vision-Language-Action Model
- SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- What Can RL Bring to VLA Generalization? An Empirical Study
- Vision-Language Models as Success Detectors
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- DextrAH-RGB: Visuomotor Policies to Grasp Anything with Dexterous Hands
- VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation
- ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving