SimVLA: Zero-Shot Sim-to-Real VLA Learning for Mobile Manipulation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "SimVLA: Zero-Shot Sim-to-Real VLA Learning for Mobile Manipulation".
Rosa: The gist SimVLA introduces an end-to-end framework that trains Vision-Language Models for mobile manipulation entirely on synthetic simulation data without requiring teleoperation, demonstrating zero-shot transfer to real home environments.
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: So we've covered the basic idea of SimVLA, which is this end-to-end framework designed to train Vision-Language Models for mobile manipulation entirely within simulation and then test if they work in the real world without any direct teleoperation.
Dev: The thesis here is that while large datasets have helped LLMs and VLMs, robot action data remains too expensive and limited, especially for long tasks where the robot needs to see beyond its current view or has a complex action space. SimVLA claims simulation offers a scalable alternative to this data bottleneck for training these models.
Taro: So they are essentially arguing that you can substitute the need for massive amounts of real-world interaction data with high-quality, diverse synthetic data generated through structured methods like composing atomic skills and using simulator state supervision.
Rosa: Right. They achieve this by pre-training on SimAction and SimVQA, then post-training on a mixture including SimDeploy to get that better generalization for deployment scenarios. The key is that they use both action supervision and visual-language feedback from the simulation environment itself during the initial learning phase.
Dev: To make it concrete, they construct these kitchen scenes procedurally using a Scene Synthesizer, generating five common kitchen types with randomized furniture and materials to ensure diversity across their one hundred houses <ref:2610.11248#pg3>. Then SimAction generates about 200K trajectories spanning thirty-five tasks across all those houses.
Taro: And SimVQA provides that extra layer of supervision by using the simulation ground truth to give the model spatial, geometric, and subtask-level feedback, which is crucial for understanding *where* things are in the scene beyond just knowing *how* to grab them.
Rosa: So it’s a layered approach. They use SimAction for action supervision and SimVQA for visual-language supervision simultaneously during pre-training. This joint optimization of continuous action prediction, tokenized action prediction, and that VQA supervision in a single forward pass is what drives the initial learning process.
Dev: That dual supervision seems like it’s what gives the model the necessary grounding to translate visual language instructions into precise robot movements effectively before they even touch a real robot.
Taro: I think that addresses one of the naive questions people would ask: how does it handle ambiguity? If the world misbehaves, what happens when things aren't exactly as expected in the simulation? The idea is that by using deployment rollouts, they prepare it for those unpredictable scenarios.
Rosa: Exactly. The SimDeploy data substantially improves specific tasks like putting a mug in a sink or a bowl in a drawer by adding deployment trajectories with diverse layouts and retry behaviors, which makes the model much more robust when it encounters variations outside the initial training set.
Dev: That post-training phase on SimAction and SimDeploy is where they really test generalization. It's not just learning to do one thing; it's learning to handle a whole distribution of deployment scenarios in simulation before we let it touch the real Anubis robot.
Taro: So the main claim here is that this combination—using action data, visual-language supervision, and deployment rollouts in simulation—can enable zero-shot sim-to-real mobile manipulation for these complex robotic tasks.
Rosa: That's the big takeaway for now: SimVLA provides a complete pipeline where simulation generates not just the actions but also the necessary supervisory signals to train a VLA that is ready to perform those tasks in reality without needing real-world training data collection.
Conclusion: Dev: So wrapping up on "SimVLA: Zero-Shot Sim-to-Real VLA Learning for Mobile Manipulation," the authors are proposing this framework where simulation acts as a source for action trajectories, visual language supervision, and deployment-oriented rollout data to enable zero-shot sim-to-real learning.
Rosa: The implication is that we can move towards building mobile manipulation AI systems that are more generalizable because they don't rely on collecting massive amounts of specific, expensive real-world interaction data for every single robot.
Taro: It suggests a path where simulation doesn't just generate trajectories; it becomes a source of supervision and rollout data, which is critical for making VLA learning applicable to the real world without constant human intervention.
Dev: Essentially, SimVLA enables more generalizable mobile manipulation by proving that combining action, visual-language, and deployment-oriented supervision within simulation can achieve zero-shot transfer from simulated training to unseen real environments.
Rosa: So it means we have a powerful tool for building more robust mobile manipulation AI by leveraging simulation not just for generating the raw motions but also as a source of structured learning signals.
Taro: It’s about making sure that when these models operate in the real world, they have the spatial reasoning capabilities to handle unexpected situations because they were trained on a wide variety of simulated conditions.
Dev: That's the big picture for this paper: simulation can be used not just as a data generator for actions, but as a source of visual-language supervision and deployment-oriented rollout data for zero-shot sim-to real VLA learning in mobile manipulation.
Kyoungin Baik, Youngwoon Lee
Yonsei University · Seoul National University
cs.RO
Submitted: 2026-10-08
Updated: 2026-10-08
Project page: https://kyounginbaik.github.io/simvla
The gist: The gist SimVLA introduces an end-to-end framework that trains Vision-Language Models for mobile manipulation entirely on synthetic simulation data without requiring teleoperation, demonstrating
Key concepts
- SimAction
- This is a large dataset of approximately 200K robot actions generated in simulation. It is created by combining basic skills into complex tasks and planning collision-free paths. It provides the model with the necessary action supervision for learning how to move robots in diverse manipulation scenarios.
- SimVQA
- This dataset uses privileged simulator state to give visual-language supervision. Instead of just showing what actions were taken, SimVQA provides spatial and geometric information about the scene. This helps the VLA model develop strong spatial reasoning skills before it even starts learning to perform tasks.
- SimDeploy
- This data consists of policy rollouts in varied simulated environments, exposing the model to deployment-oriented trajectories. It includes diverse layouts, viewpoints, and retry behaviors. Post-training on this data is critical for making the learned policies robust enough to handle unseen real-world situations.
- Hybrid System Identification
- This mechanism calibrates physical parameters of the real robot by minimizing errors between simulation and reality across different metrics (position, rotation, joint space). This helps bridge the gap between simulated physics and actual robot performance, leading to better generalization in the real world.
Terminology
Summary
The gist SimVLA introduces an end-to-end framework that trains Vision-Language Models for mobile manipulation entirely on synthetic simulation data without requiring teleoperation, demonstrating zero-shot transfer to real home environments.
SimVLA Framework and Data Sources
SimVLA is a zero-shot sim-to-real framework that trains VLAs entirely from simulation data for mobile manipulation SimVLA is first pre-trained on two complementary simulation-derived datasets: SimAction, a large-scale robot action dataset spanning 35 diverse mobile manipulation tasks, generated by composing atomic skills, and SimVQA, which leverages privileged simulator state to provide spatial, geometric, and subtask-level visual-language supervision We further post-train SimVLA on a mixture of SimAction and SimDeploy Together these data sources allow SimVLA to exploit simulation beyond action generation alone.
Data Generation Components
The framework constructs procedurally generated kitchen scenes using Scene Synthesizer to instantiate five common kitchen types, randomizing furniture heights, object arrangements, cabinet configurations, and materials SimAction is a large-scale synthetic action dataset comprising approximately 200K trajectories spanning 35 mobile manipulation tasks across 100 procedurally generated houses. This data is generated without teleoperation by composing atomic skills into task programs and converting their subtask goals into collision-free robot trajectories through motion planning. SimVQA leverages simulation ground-truth to provide spatial, geometric, and subtask-level supervision for VLA co-training SimDeploy consists of policy rollouts in diverse simulated environments, exposing the model to deployment-oriented trajectories with broader scene layouts.
Two-Stage Training Pipeline
SimVLA is trained in two stages The model is first pre-trained on SimAction for action supervision and SimVQA for visual-language supervision. We then roll out the policy in diverse simulated environments to collect SimDeploy, and post-train SimVLA on a mixture of SimAction and SimDeploy, which we find critical for generalization to unseen real-world environments. During pre-training, SimVLA jointly optimizes continuous action prediction, tokenized action prediction, and SimVQA supervision in a single forward pass.
Evaluation and Results
We evaluate SimVLA on long-horizon mobile-manipulation tasks in the real world, including mock kitchens, a pantry, and a real home SimVLA outperforms policies trained on 50 in-domain real-world demonstrations in our evaluation settings. It also maintains robust performance in a pantry and a real home, where both baselines degrade sharply SimVLA achieves 54.4% average task progress, outperforming the pi0.5 policy fine-tuned on 50 indomain real-world demonstrations in our evaluated settings.
Key Contributions to Robustness
The framework incorporates several mechanisms to enhance robustness and generalization Hybrid system identification calibrates physical parameters by jointly minimizing end-effector position error, end-effector rotation error, and joint-space MSE between simulation and the real robot. Extensive domain randomization across six categories—including camera perturbations, kitchen scene properties, object physical and visual properties, robot dynamics, and lighting conditions—exposes the policy to a broad distribution of physical and visual conditions during training. This suggests that large-scale simulation can mitigate the limitation of limited trajectory coverage in real-world demonstrations for OOD generalization.
Conclusion on SimVLA's Efficacy
SimVLA demonstrates that combining action, visual-language, and deployment-oriented supervision in simulation can enable zero-shot sim-to-real mobile manipulation. The results suggest that combining action, visual-language, and deployment-oriented supervision in simulation can enable zero-shot sim-to real mobile manipulation. SimVLA enables more generalizable mobile manipulation. The paper concludes that simulation can serve not only as a source of action trajectories, but also as a source of visual-language supervision and deployment-oriented rollout data for zero-shot sim-to real VLA learning in mobile manipulation.
Limitations and Future Work
Although SimVLA demonstrates zero-shot sim-to real transfer for mobile manipulation, it has several limitations First, we do not isolate the contributions of individual domain randomization factors. Second, our tasks focus on rigid-object manipulation in kitchen-like environments; extending to deformable objects, dexterous manipulation, and broader household settings remains future work. Third, since SimVLA is a simulation-based data generation framework rather than a cross-embodiment transfer method or a new VLA architecture, our evaluation is limited to a subset of embodiments and architectures. Finally, SimAction generation relies on human-defined skill APIs; future work could leverage LLMs to expand skills as code-based policies, enabling a more autonomous pipeline. The paper also notes that the hybrid identification objective is best on both metrics but gains only 1.4 mm in mean error over joint-space identification and is within noise of EEFspace identification, suggesting that the primary benefit comes from performing system identification itself. The SimVQA pre-training provides the model with accurate spatial reasoning capabilities beyond action prediction. The SimDeploy data substantially improves put mug / bowl / bottle in sink and put bowl in drawer as it adds deployment trajectories with diverse layouts, viewpoints, and retry behaviors.
Experimental Setup Details
The real-world experiments use the Anubis robot, a bimanual mobile manipulator with two 6-DoF arms and a 3-wheel omnidirectional base The policy action space is 23-D, comprising a 10-D end-effector representation for each arm (3-D position, 6-D rotation and 1-D gripper), together with the base velocities. The evaluation setup includes testing in three real-world environments: a mock kitchen, a pantry, and a real home. The task progress descriptions provide finergrained evaluation beyond binary success metrics for each task. Table 3 shows that hybrid identification cuts mean EEF error by 6.3× (83.4 → 13.3 mm) and maximum error by 3.0× (120.7 → 40.6 mm) over default gains. The SimVQA pre-training provides the model with accurate spatial reasoning capabilities beyond action prediction. The SimDeploy data substantially improves put mug / bowl / bottle in sink and put bowl in drawer as it adds deployment trajectories with diverse layouts, viewpoints, and retry behaviors. The paper also notes that the SimVQA pre-training provides the model with accurate spatial reasoning capabilities beyond action prediction. The SimDeploy data substantially improves put mug / bowl / bottle in sink and put bowl in drawer as it adds deployment trajectories with diverse layouts, viewpoints, and retry behaviors. The paper also notes that the SimVQA pre-training provides the model with accurate spatial reasoning capabilities beyond action prediction. The SimDeploy data substantially improves put mug / bowl / bottle in sink and put bowl in drawer as it adds deployment trajectories with diverse layouts, viewpoints, and retry behaviors. The paper also notes that the SimVQA pre-training provides the model with accurate spatial reasoning capabilities beyond action prediction <ref:
Improvements for AI systems
-
Bold header: Knowledge-isolated co-training during pre-training. This allows SimVLA to
jointly optimize continuous action prediction, tokenized action prediction, and SimVQA supervision
while preventingunstable gradients from the randomly initialized action head from degrading the pre-trained VLM backbone,
ensuring a robust initial representation. -
Bold header: Hybrid system identification for dynamic calibration. The procedure
jointly minimizing end-effector position error, end-effector rotation error, and joint-space MSE between simulation and the real robot
to achieve asubstantial effect
in reducing mean EEF error by 6.3× compared to default gains when calibrating physical parameters. -
Bold header: Domain randomization for OOD robustness. Applying extensive domain randomization across six categories
exposes the policy to a broad distribution of physical and visual conditions,
which makes SimVLArobust performance
in OOD environments, as shown by its superior task progress compared to baselines underOOD-pos
andOOD-vis.
-
Bold header: Deployment-oriented post-training for generalization. Post-training on a mixture of SimAction and SimDeploy allows the model to benefit from
policy rollouts across diverse simulated environments,
which is found to improverobustness and generalization to unseen real-world environments.
-
Bold header: Fine-grained task progress evaluation. Defining discrete task progress labels (e.g., Task 2:
Failed to open the drawer (e.g., slipped during opening)
) providesfiner grained evaluation beyond binary success metrics
for complex mobile manipulation sequences like opening and closing drawers.
Sources
- PaliGemma: A versatile 3B VLM for transfer
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- PaLI: A Jointly-Scaled Multilingual Language-Image Model
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
- MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
- MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation
- Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- DextrAH-RGB: Visuomotor Policies to Grasp Anything with Dexterous Hands
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving