Getting Out and Getting Back: World and Behavior Grounding in Real2Sim2Real Co-Training
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Getting Out and Getting Back".
Dev: Simulation can expand scarce real demonstrations for co-training, yet how world fidelity and similarity to human behavior affect policy performance remains unclear.
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: So we're talking about this paper, "Getting Out and Getting Back: World and Behavior Grounding in Real2Sim2Real Co-Training," which basically looks at how simulation can help us train policies using real demonstrations. The authors are trying to figure out what really matters when you're mixing simulated data with real experience, specifically focusing on world fidelity and how closely the simulated actions mimic human motion.
Dev: It sounds like they're tackling that tricky gap between simulation and reality, which is always a concern for anyone working on robotics where the loop rate and latency are critical issues. What's the central idea they put forward in this paper regarding world grounding versus behavior grounding?
Taro: The core hypothesis seems to be that world grounding is the main factor because it dictates whether actions learned in simulation actually have any effect on the real robot, while behavior grounding keeps the simulated stuff looking more like what a human would actually do when things get messy.
Rosa: Exactly, they set up this real2sim2real pipeline to test those axes separately by varying both world grounding and behavior grounding independently. It matters because they are testing how these two factors interact when you’re co-training policies from scratch or post-training foundation models using one hundred real teleoperation demonstrations <ref:2610.00821#pg0>.
Dev: That experimental setup is pretty thorough, varying the configurations to cover all four possibilities: grounded world/ungrounded behavior, ungrounded world/grounded behavior, ungrounded world/ungrounded behavior, and finally grounded world/grounded behavior. I'm curious about how those specific combinations led to the performance metrics they found.
Taro: The results show that fully grounded co-training can raise success rates on a dynamic dexterous pick-and-sort task from fifty-two percent up to eighty-six percent <ref:2610.00821#pg0>. More specifically, world grounding alone improved success by eighteen percentage points, and behavior grounding helped by ten percentage points when averaged across all the configurations.
Rosa: That jump from fifty-two percent to eighty-six percent is a significant performance gain, showing that world grounding allows policies to actually use simulated experience beyond what's available in the real data alone <ref:2610.00821#pg0>. But they also point out that behavior grounding becomes more important when the world grounding isn't perfect.
Paper summary: Dev: I'm interested in the qualitative observations because from an engineering standpoint, knowing *why* a policy switches its strategy is crucial for diagnosing failure modes during deployment. The paper notes that policies switch between using real-like approaches in states they cover and switching to simulated behavior when they miss something, even if that simulated motion looks quite different from human motion.
Taro: That switching mechanism suggests the policy is using the simulation data as a fallback or a correction tool when it encounters situations outside its direct training coverage, which is important for autonomy in unpredictable environments. Furthermore, failures like dropping objects or drifting seem most common in policies trained on data that lacked both accurate dynamics and human-like behavior to rely on.
Rosa: That paints a picture of how the system behaves when it's struggling; it's not just about having one good source of data, but having the right balance between physics fidelity and motion similarity. This whole investigation into world grounding versus behavior grounding really helps clarify how we can safely use simulation to prepare policies for real-world deployment.
Dev: Thinking about the loop rate and latency, if a policy is relying heavily on simulated behavior for corrections in those missed states, we need to make sure that switch happens fast enough and smoothly so it doesn't introduce noticeable jitter or instability when running on physical hardware. How does this grounding distinction affect the practical deployment timeline?
Taro: The authors imply that world grounding is what lets the policy actually operate on the real robot successfully in terms of action effect, which is a big deal for deployment feasibility, but behavior grounding determines how well it can recover its intended motion once it's operating.
Rosa: So, to summarize this paper, "Getting Out and Getting Back: World and Behavior Grounding in Real2Sim2Real Co-Training," the main point is that world grounding determines if simulated actions work on the real robot, while behavior grounding dictates whether the policy can return to what a human would actually do.
Paper summary: Dev: And considering those findings, what does this mean for us when we're deploying these systems outside of a highly controlled lab setting? Can we expect this level of performance and reliability in an open environment?
Taro: The implication is that as long as the world grounding is solid, you get a strong base from simulation, but if the real world throws something completely unexpected at it, the behavior grounding gives the policy some mechanism to try and revert to something more sensible.
Rosa: So we're looking at a complementary relationship where world grounding provides the foundation for using simulated experience beyond what's physically possible in real demonstrations, and behavior grounding acts as a safety net for matching human intent when things go awry.
Dev: That makes sense from an engineering standpoint; we need both reliable physics modeling and robust motion matching to handle the inevitable discrepancies that arise between the simulation and the physical system during long-running tasks.
Taro: I think the real world will test this switching mechanism constantly, especially in long-horizon tasks where recovering demonstrated behavior becomes even more critical for success.
Rosa: And looking ahead, this suggests that grounded simulation remains a beneficial tool when we are co-training foundation models because it helps compensate for the scarcity of real demonstrations.
Dev: It seems like the authors suggest that the utility of this approach extends beyond just training policies from scratch; it still helps post-trained foundation models achieve performance levels comparable to training on a larger set of real demonstrations.
Taro: I think the future work mentioned points toward exploring longer-horizon tasks where getting back to demonstrated behavior should be a more central focus for their research efforts.
Rosa: So, the paper "Getting Out and Getting Back: World and Behavior Grounding in Real2Sim2Real Co-Training" provides a framework showing that separating world grounding from behavior grounding helps us understand how simulation data can be leveraged effectively in real-world co-training scenarios.
Conclusion: Rosa: So we've been looking at how this paper dissects world grounding and behavior grounding in real2sim2real co-training, and now it’s time to talk about the title and those authors.
Dev: I’m ready to hear what they actually found regarding the implications of separating those two concepts.
Taro: I'm curious how this framework changes our view on training policies that are supposed to operate in unpredictable environments.
Rosa: Essentially, it shows us that world grounding is about whether the simulation can actually translate into physical action, and behavior grounding is about whether the policy can recover human-like movements when things go wrong.
Dev: That distinction between those two axes seems really important for understanding system reliability under stress.
Taro: And I think that ability to switch between simulated and demonstrated behavior during a rollout is where the real autonomy potential lies, especially if the world misbehaves unexpectedly.
Rosa: Exactly, this suggests that for deployment outside of a perfectly controlled lab, we need to consider both how well the simulation mimics physics and how well it mimics human intent.
Dev: From an engineering standpoint, that means we have to design our systems with mechanisms that can handle those switching moments without introducing instability or latency spikes in the loop rate.
Taro: I agree; if a policy can reliably fall back on real demonstrations when the simulated path fails, that opens up possibilities for more robust real-world operation.
Rosa: It really frames the challenge as needing both high-fidelity physics modeling and strong behavioral alignment to bridge the sim2real gap effectively.
Dev: So, where do we go from here with this understanding of grounding? What does this mean for long-term deployment scenarios?
University of Cambridge
cs.RO
Submitted: 2026-09-30
Updated: 2026-10-02
Comments: Project Website: https://industrialnext.github.io/r2s2r-grounding/
Code: https://github.com/google-deepmind/mujoco_warp
Project page: https://graspgenx.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 87/100
The gist: Simulation can expand scarce real demonstrations for co-training, yet how world fidelity and similarity to human behavior affect policy performance remains unclear.
Key concepts
- World Grounding
- This refers to how accurately the simulator matches the real robot's physical environment, including rendering, physics, and control differences. It ensures that actions learned in simulation have a meaningful effect on the actual robot.
- Behavior Grounding
- This focuses on aligning simulated movement trajectories with human motion patterns. It is important when world grounding is imperfect because it allows the policy to fall back onto realistic human behavior during rollouts.
- Policy Switching Mechanism
- Policies learn to use both data sources dynamically. They adopt simulated movements in states where simulation provides coverage, even if those movements differ from human demonstrations. This switching helps the policy navigate complex scenarios.
- Foundation Model Benefit
- Grounding simulation data still helps foundation models after they are pre-trained on real data. This suggests that high-quality, grounded simulated experience can compensate for gaps in real-world data coverage.
Terminology
Summary
Simulation can expand scarce real demonstrations for co-training, yet how world fidelity and similarity to human behavior affect policy performance remains unclear. The gist: Grounded simulation remains beneficial when co-training foundation models.
The Core Problem and Hypothesis
The paper investigates the downstream effect of the sim2real gap by separating the grounding of simulation data into two axes: World Grounding and Behavior Grounding. World grounding aligns the simulator with the real system through rendering, physics, and control differences, while behavior grounding aligns simulated trajectories with human motion. The authors hypothesize that world grounding is the primary factor in co-trained policy performance since it determines whether actions learned in simulation have the same effect on the real robot.
However, because world grounding is never perfect in practice, behavior grounding plays an important role in keeping simulated behavior close to real demonstrations, which the policy can fall back on.
The Real2Sim2Real Co-Training Pipeline
To test these axes independently, a real2sim2real pipeline
was built that varies world grounding and behavior grounding. The setup involves generating data under four configurations: Grounded World/Ungrounded Behavior, Ungrounded World/Grounded Behavior, Ungrounded World/Ungrounded Behavior, and Grounded World/Grounded Behavior. Each of these simulated datasets is then combined with 100 real teleoperation demonstrations to co-train policies from scratch or post-train foundation models.
The evaluation involves testing policies over 50 real-world rollouts that include both nominal and shifted conditions.
Experimental Results on Performance Gains
The study measured the impact of these grounding configurations on success rates for a dynamic dexterous pick-and-sort task. When comparing the four configurations, fully grounded co-training raises success from 52% to 86%; averaged across configurations, world grounding improves success by 18 percentage points and behavior grounding by 10.
This demonstrates that world grounding lets policies use simulated experience beyond real-data coverage,
while behavior grounding matters mainly when world grounding is imperfect.
For foundation models, the paper shows that with 100 real demonstrations and 1,500 fully grounded simulated trajectories,
the model reaches 90%, which matches or exceeds fine-tuning on 1,000 real demonstrations.
Qualitative Observations on Policy Behavior
Qualitative analysis revealed how policies utilize both data sources. Policies switch between the two within a rollout, adopting simulated behavior in states covered only by simulation, even when it differs substantially from human behavior.
This switching is observed in rollouts where the robot began with a real-like approach but missed the object, then switched to visibly sim-like motion,
suggesting reliance on simulated data for corrections. Furthermore, failures in releasing objects or drifting occurred mainly in policies trained on world- and behavior-ungrounded data,
which lacked both accurate dynamics and human-like behavior to fall back on.
Conclusion and Future Directions
The study concludes that world grounding determines whether borrowed simulation behavior works on the real robot, while behavior grounding dictates the policy's ability to return to demonstrated behavior. The findings suggest that grounded simulation remains beneficial when co-training foundation models,
extending the utility of this approach beyond policies trained from scratch. Limitations noted include the qualitative nature of latent-space analysis and the difficulty in definitively attributing actions solely to real or simulated data within embeddings. Future work is suggested for exploring longer-horizon tasks where returning to demonstrated behavior should be more critical.
Key Grounding Definitions
(a) World Grounding:
Grounded World:
-
We calibrate camera geometry and blur, randomize lighting and object color, register robot and conveyor geometry, and use CAD objects with measured masses.
-
For Ungrounded World, we use ideal cameras, tape-measured poses, a supplied cell layout, nominal kinematics, and calculated controller gains.
(b) Behavior Grounding:
-
For Grounded Behavior, we adopt MimicGen [7] to shift teleoperated demonstrations to new object poses, sizes, and belt speeds while preserving their motion.
-
For Ungrounded Behavior, we synthesize trajectories without human motion references using GraspGen-X proposals adapted to hand closure geometry.
(c) Policy Switching Mechanism:
-
Policies
imitates the real demonstrations in states they cover and relies on simulated behavior elsewhere.
-
The policy's ability to return to real behavior is dependent on
behavior grounding whether the policy can return to demonstrated behavior.
(d) Foundation Model Benefit:
-
Grounded simulation still benefits post-trained foundation models, suggesting that
pre-training partly compensates for ungrounded data.
-
Co-training with grounded simulation allows the model to reach performance levels matching or exceeding those achieved by training on a larger set of real demonstrations.
(e) Performance Comparison:
Improvements for AI systems
As a fastidious researcher, I have analyzed this paper, Getting Out and Getting Back: World and Behavior Grounding,
which focuses on improving robot manipulation policies through a novel real-to-simulation-to-real (Real2Sim2Real) co-training pipeline.
The core finding is that separating simulation grounding (how well the simulation matches reality, i.e., world grounding) and behavior grounding (how closely simulated motion resembles human demonstrations) allows for a controlled study of their complementary roles. The key takeaway is that world grounding is crucial for generalization to unseen conditions, while behavior grounding acts as a fallback when the world fidelity is imperfect.
Here are the specific improvements and capabilities this research enables in AI systems:
-
Improve Robustness and Generalization in Real-World Robotics:
-
Enhance Policy Reliability under Novel Conditions (Extrapolation):
-
Enable Efficient Data Utilization with Scarce Real-World Demonstrations:
-
Facilitate the Development of Hybrid, Adaptive Robotic Policies (Mixture of Learned Behaviors).
5.1 Specific Improvements and Capabilities:
-
[] The AI system can achieve significantly higher task success rates (e.g., moving from 52% to 86% in the pick-and-sort task) when leveraging a combination of real demonstrations and simulation data, specifically by ensuring the simulated environment accurately reflects real physics and dynamics (World Grounding).
-
[] The policy can effectively generalize to conditions absent from the training data (e.g., faster belt speeds or novel object placements) because world grounding ensures that actions learned in simulation have the same physical effect on the real robot, even when those conditions are unseen in the real demonstrations.
-
[] The system can utilize a mixture of policies: it can imitate expert human behavior precisely in states covered by real-world data and seamlessly switch to relying on simulated behavior for states where simulation is available but doesn't perfectly match reality (Hybrid Behavior).
-
[] For Foundation Models, the system can achieve state-of-the-art performance even when starting with a limited number of real demonstrations (100 vs. 1,000) by combining them with high-fidelity, world-grounded simulation data. This makes high-quality robotic skills achievable even under severe data scarcity.
-
[] The system can be designed to recover from simulated behavioral errors or unreliability (e.g., failing to release an object in simulation) by falling back on human-like, behavior-grounded trajectories when the world fidelity is low, ensuring the robot returns to a
real-like
state. -
[] The system can be optimized using targeted grounding strategies: prioritizing World Grounding for broad robustness across different physical scenarios and utilizing Behavior Grounding specifically to bridge the gap when world fidelity is imperfect.
In summary, this research enables the creation of more reliable, robust, and data-efficient robotic AI systems capable of operating effectively in complex, real-world environments by intelligently managing the discrepancy between simulation and reality through controlled grounding mechanisms.
Sources
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
- What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World
- Empirical Analysis of Sim-and-Real Cotraining of Diffusion Policies for Planar Pushing from Pixels
- Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation
- Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation
- Video2Sim2Real: Full-Stack Autonomous Dexterous Skill Acquisition from a Single Human Video
- A Systematic Study of Data Modalities and Strategies for Co-training Large Behavior Models for Robot Manipulation
- Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models
- VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation
- Sim-to-Real: Learning Agile Locomotion For Quadruped Robots
- Grounding Sim-to-Real Generalization in Robotic Manipulation: An Empirical Study with Vision-Language-Action Models
- DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning
- Beyond Human Demonstrations: Diffusion-Based Reinforcement Learning to Generate Data for VLA Training
- RDGen: Demonstration Generation for High-Quality Robot Learning via Reinforcement Learning
- Imitating Task and Motion Planning with Visuomotor Transformers
- GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
- SAM 2: Segment Anything in Images and Videos
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving