Measuring Asset and Scene Reconstruction Effects in Real-to-Sim Robot Evaluation

arXiv:2610.00731 · cs.RO · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "Measuring Asset and Scene Reconstruction Effects in Real-to-Sim Robot Evaluation".

Rosa: Simulated evaluation is increasingly used alongside real-world evaluation of robot policies because it is cheaper and easier to repeat; however,

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: So, we're diving into this paper titled "Measuring Asset and Scene Reconstruction Effects in Real-to-Sim Robot Evaluation," which is really looking at how much the way we build our digital worlds impacts how well an AI policy actually performs when it hits the real robot.

Dev: That's right, Rosa, and it tackles a crucial question about using simulation for training versus deploying in the actual physical world because if the simulation doesn't match reality closely enough, all that training just won't translate well to things happening on a real robot.

Taro: I think what they are focusing on is that simply having a policy trained in simulation isn't enough if the digital environment itself is fundamentally different from what the robot encounters in reality.

Rosa: Exactly, and according to this paper, they set up a comparison between two different ways of building their simulated robot cell: an "authored reconstruction" and a "default reconstruction."

Dev: The authors are testing whether improving that scene fidelity—by using better geometry estimates and authored physics—actually closes the gap between what happens in simulation and what happens on the actual hardware.

Taro: And I wonder if this is just about making things look pretty in simulation, or if it's fundamentally changing how the robot interacts with objects when physics are different.

Rosa: Well, according to their summary of "Measuring Asset and Scene Reconstruction Effects in Real-to-Sim Robot Evaluation," the main takeaway is that refining the scene reconstruction using high visual fidelity and authored physics makes the simulation much more faithful to what happens in the real world, which effectively narrows that sim-to-real gap.

Dev: They constructed two versions of a bimanual robot cell: one where they used object geometry at estimated metric scale, projected textures, authored physics, and their own scene splat—that's the authored reconstruction—and another one based on an open-source recipe with engine-default physics and a Gaussian-splat scene.

Taro: So they are explicitly separating the variables to see which components—geometry, texture projection, or especially the physics modeling—make the biggest difference in bridging that gap.

Title and authors: Rosa: Right, and they used the same object photographs and scene video inputs for both versions while keeping everything else like the robot model and controller fixed between them so they could isolate what was changing.

Dev: The experimental setup involved ten "cells," each run twenty times in each reconstruction, evaluating five tasks with two policies, pi zero point five and MolmoAct2, which gave them ten task-policy pairs overall.

Taro: That's a pretty rigorous setup for testing the robustness of these different simulation environments across various complex manipulation scenarios.

Rosa: They graded the real trials using Robocurve, which is an evaluator-operated robot running its own protocol with specific scoring rules like "grasped and lifted" or "in bin on its side," which gives them a very concrete measure of success.

Dev: When looking at the comparison metrics for this paper, they focused on four main areas: score error, Pearson correlation, progress disagreement, and failure-stage disagreement across those ten simulated and ten real cell means.

Taro: I'm interested in what they found about the correlation measure; does it mean that if a policy scores high in simulation, it actually performs well when we test it on the real robot?

Rosa: The results showed that the authored reconstruction achieved a score error of six point nine seven percentage points compared to seventeen point five four for the default reconstruction, which is a reduction of about ten point five six percentage points in that measure.

Dev: That score error reduction is significant, and they also saw an improvement in progress disagreement by eight point three two percentage points when comparing the authored version to the default one across those five tasks and two policies <ref:2610.00731#pg0>.

Taro: It sounds like the authors are claiming that this combined approach of metrically scaled geometry, authored physics, and their own scene reconstruction is what really closes that gap relative to what's available in the literature.

Rosa: That's precisely what they concluded; they found that a reconstruction with metrically scaled geometry, authored physics and our own scene reconstruction taken together narrows the sim-to-real gap compared to the default recipe used in previous work.

Dev: They further noted that while correlation and failure-stage disagreement didn't strictly separate the two reconstructions when looking at ten cells, they found that the authored reconstruction was consistent with reality in nine out of ten cells, compared to only five out of ten for the default version <ref:2610.00731#pg0>.

Title and authors: Taro: It’s interesting that they also identified a shared cause for performance gaps, pointing to the control gap where the simulated arm and policy and its controller are identical in both reconstructions.

Rosa: So, even with better reconstruction assets, if the control loop itself has a mismatch between simulation and reality, that's still causing some of the performance differences they observed.

Dev: I agree; it suggests that while asset quality is important, we can't ignore the underlying control discrepancies when evaluating policies.

Taro: Looking ahead, what does this mean for deploying these policies in more complex, messy real-world situations where things don't go exactly as planned?

Rosa: The implication is that if we invest time and effort into creating high-fidelity assets with proper physics rather than relying on defaults, the simulation becomes a much more reliable testing ground for the robot's capabilities.

Dev: From an engineering standpoint, it means we can trust the simulated performance metrics to predict real-world success with much higher confidence as long as we follow this enhanced reconstruction pipeline.

Taro: For autonomy research, this suggests that future VLA policies could be evaluated against environments that actually mimic the physical constraints and visual noise of reality, rather than idealized simulations.

Rosa: So, to wrap up on "Measuring Asset and Scene Reconstruction Effects in Real-to-Sim Robot Evaluation," the paper strongly supports building our digital twins with metrically scaled geometry and authored physics instead of just using open-source defaults.

Dev: It's a solid methodological improvement because it directly measures the impact of those specific reconstruction choices on task success metrics like score error.

Taro: I think the real world impact here is that we can start designing robots for deployment environments that are much closer to their true operational context, which is where autonomy really needs to mature.

Rosa: Well said, Taro, and Dev, this work gives us a clear direction on how to make simulation serve as a better predictor of real-world performance by focusing on the fidelity of the physical representation itself.

The paper's summary: Rosa: So, to recap, this paper is really showing how much better our simulation results are when we invest time in making the digital environment match reality through detailed object modeling and scene reconstruction rather than just using off-the-shelf tools.

Dev: That’s right, Rosa; it’s about moving away from those default recipes that often introduce a lot of unseen discrepancies between the simulated and real robot experiences.

Taro: I think the core idea here is that if you want to test an AI policy on a physical robot, you need to ensure the digital twin it's learning in is as physically accurate as possible, which this research directly addresses.

Rosa: Exactly; they’re demonstrating that by using metrically scaled geometry and physics authored specifically for their scene, they significantly reduce the gap between simulated performance and real-world outcomes.

Dev: From an engineering standpoint, that means we can have much higher confidence when we look at score errors and progress disagreement in simulation because those numbers will track reality much more closely.

Taro: And when things go wrong in the real world, this kind of fidelity should give us a better chance of predicting *why* it failed or succeeded in the first place, which is critical for building truly robust autonomy.

Rosa: It really opens up possibilities for how we develop these policies; instead of just training them in whatever simulation assets are easiest to get, we’re encouraged to prioritize creating high-fidelity digital environments that accurately represent material properties and physical constraints.

Dev: That fidelity directly impacts the control loop's stability in the simulation, so a better reconstruction means fewer false positives or misleading feedback signals that could skew policy training.

Taro: I wonder how this will help when we move from controlled lab settings to genuinely chaotic, unpredictable real-world scenarios where things don't follow the clean physics they see in a dataset.

Rosa: That’s the big question; it suggests that a more accurate simulation foundation is one of the necessary steps toward deploying AI systems in truly messy, unstructured environments where we can't just assume perfect world modeling.

The paper's improvements: Taro: So, to wrap up on the technical side, the paper isn't just presenting results; they’re proposing specific ways to build better simulations using "authored" reconstructions instead of relying on generic defaults.

Rosa: That makes sense; they aren't just saying the default way is bad, they’re giving us a blueprint for building a much more faithful digital twin by focusing on metric scale and authored physics.

Dev: And those suggested improvements are really focused on making sure the geometry isn't just a pretty mesh but something that respects real-world physical constraints, like using Signed Distance Fields for better collision modeling.

Rosa: Right, and they’re pushing for a unified pipeline where object reconstruction, scene splatting, and physics are all built together from the start to avoid those mismatched components we saw before.

Dev: I agree; having that single authored pipeline should help us manage latency and failure modes more predictably because the entire system is designed around consistent physical parameters.

Taro: It’s interesting how this ties in with other work, like FlashDexRetarget, which focuses on generating high-success motion data; better scene fidelity could also improve how well those generated motions translate into real-world performance.

Rosa: Exactly; if the visual representation is spot on, then the learned behaviors should be more transferable to deployment outside the lab, and that’s what I’m most interested in as a field roboticist.

Dev: From my side, it means we can test policy robustness against physics failures earlier in development because we're using models that are actually grounded in physical reality rather than guesswork.

Taro: What about the limitations they pointed out? They admitted that because they’re changing so many things—geometry, scale, physics—it’s hard to isolate which single factor causes the biggest improvement yet.

Rosa: That means future work needs to break it down further; testing each of those factors individually will help us understand exactly where we can spend our effort to get the best sim-to-real transfer.

Conclusion: Rosa: So we've covered how this paper, "Measuring Asset and Scene Reconstruction Effects in Real-to-Sim Robot Evaluation," shows that investing in high-fidelity scene reconstruction dramatically narrows the gap between simulation and real robot performance metrics.

Dev: That’s a solid summary, Rosa; it really boils down to making sure the digital environment is physically consistent so the control loop doesn't get confused by phantom dynamics.

Taro: I think what this implies for autonomy research is that we can start trusting simulation evaluations more when we are dealing with complex, messy real-world environments because the underlying representation is much closer to physical reality.

Rosa: Exactly; it suggests that the quality of our digital twins isn't just about having more data, but about making sure that data respects metric scale and physics authored specifically for the task at hand.

Dev: And for us engineers, it means we can focus less on tuning the simulation to match reality and more on ensuring our control loop itself is robust enough to handle those high-fidelity inputs without getting unstable.

Taro: I’m still thinking about when this works outside the lab; if these reconstructions are good enough, could we see a real-world deployment lasting for hours instead of just a few minutes of controlled testing?

Rosa: That’s the million-dollar question; if the fidelity holds up under those conditions, it means we can deploy these policies in settings that are genuinely more complex than what we can currently simulate perfectly.

Dev: I’d be interested to see how these metrics hold up when we introduce high latency or intermittent sensor failures during extended operational runs, because that's where the control loop really tests its limits.

Taro: That brings us to the next topic: how do these improved reconstruction methods affect our ability to handle unexpected world misbehavior?

Rosa: We’ll definitely be talking about that next; it’s about testing what happens when the simulation starts breaking down in ways that aren't just simple visual errors.

cs.RO

Submitted: 2026-09-30

Updated: 2026-10-02

Comments: v2: added link to the code and data repository. Code: https://github.com/Kaedim/yam-sim-harness-oss

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: Simulated evaluation is increasingly used alongside real-world evaluation of robot policies because it is cheaper and easier to repeat; however, its value depends on how closely its outcomes track

Key concepts

Authored Reconstruction
This simulated environment was built by researchers incorporating several improvements over a standard setup. It included object geometry scaled to real metric measurements, projected textures, custom physics settings created by the authors, and a scene reconstruction pipeline specifically designed for better fidelity.
Default Reconstruction
This baseline simulation used an open-source recipe. It relied on a generative single-image mesh for geometry and standard engine physics (PhysX). This version served as a comparison point against the more advanced authored reconstruction, using the same input photographs and videos.
Score Error
This metric measures the absolute difference between how well the simulated robot performed on a task compared to how well it performed in real-world trials. A lower score error indicates that simulation results are closer to real-world performance metrics.
Progress Disagreement
This concept compares how far a robot gets through a specific task in the simulation versus reality by tracking which stages of scoring rubrics are reached. A lower disagreement means the simulated robot's progression closely matches the real robot's progression.

Terminology

Summary

Simulated evaluation is increasingly used alongside real-world evaluation of robot policies because it is cheaper and easier to repeat; however, its value depends on how closely its outcomes track the real robot’s. The key finding of this study is that improving the quality of the environment reconstruction through higher visual fidelity, authored physics, and metric scale makes the simulation more faithful to the real world and narrows the sim-to-real gap.

How it works

The researchers constructed two simulated versions of a bimanual robot cell: an authored reconstruction and a default reconstruction. The authored version incorporates several improvements over the baseline, including:

  1. Object geometry at estimated metric scale, projected textures, authored physics, and the authors' own scene splat.

  2. Scene reconstruction via their own pipeline.

In contrast, the default reconstruction uses an open-source recipe consisting of a generative single-image mesh for geometry and engine-default physics (PhysX) and a Gaussian-splat scene. Both reconstructions utilize the same object photographs and scene video inputs, while all other settings—robot model, controller, camera, placement budget—are held fixed between the two versions.

Experimental Setup

The study utilized a bimanual I2RT YAM station setup to evaluate ten cells, where each cell was run twenty times in each reconstruction. The evaluation involved five tasks and two policies: π0.5 and MolmoAct2. The tasks included placing bottles in a bin, stacking bowls, stacking blocks, moving a latte cup onto a board, and clearing a table into a bin. Real trials were graded by Robocurve, an evaluator-operated robot under its own protocol using specific rubrics for scoring stages (e.g., grasped and lifted, in bin on its side, upright).

Comparison Metrics

The researchers compared the ten simulated and ten real cell means using four primary agreement measures:

  1. Score error, which is the absolute difference between simulated and real cell means, averaged over ten cells. The authored reconstruction achieved a score error of 6.97 percentage points compared to 17.54 for the default reconstruction, a reduction of 10.56 percentage points (Table 4).

  2. Pearson correlation (r), which measures whether cells that score high in reality also score high in simulation, yielding r = 0.90 for the authored reconstruction and r = 0.51 for the default (Table 4).

  3. Progress disagreement, which compares how far the robot gets through a task in simulation versus reality by comparing fractions of trials reaching specific rubric stages (Table 6). The authored reconstruction showed a reduction of 8.32 percentage points compared to the default.

  4. Failure-stage disagreement, which measures whether trials stop at the same stage, showing values of 42.00 for the authored reconstruction and 48.50 for the default (Table 4).

Key Findings

The main finding is that a reconstruction with metrically scaled geometry, authored physics and our own scene reconstruction taken together narrows the sim-to-real gap relative to the default recipe used in the literature. Specifically, this combined approach resulted in:

score error is 10.56 percentage points lower and progress disagreement by 8.32 percentage points lower across five tasks and two policies, both with 95% intervals that exclude zero.

Furthermore, the authors noted that while correlation and failure-stage disagreement did not separate the two reconstructions with ten cells, the authored reconstruction was consistent with real in 9 of 10 cells compared to only 5 of 10 for the default. The study also identified a shared cause for performance gaps: the likeliest shared cause is the control gap (Section 1): the simulated arm, policy and its controller are identical in both reconstructions.

Limitations

The study acknowledged several limitations. First, because the two reconstructions differ across multiple axes (geometry, scale, physics, scene), it cannot definitively attribute the reduction solely to one factor; future work should test each factor individually. Second, a single cell (Stacking blocks with MolmoAct2) failed under both reconstructions over-scoring compared to real scores. Third, grading discrepancies between real and simulated trials by different graders affect both reconstructions equally. Finally, the study is small in scope (five tasks by two policies) and the difference in object placement variation between simulation and reality remains a factor for some gaps.

Conclusion

The authors conclude that Simulated evaluation tracks the real robot more closely when the objects and scene are built with metrically scaled geometry and authored physics than when they are built with the default open-source recipe. They release their harness, per-trial scores, configuration logs, and assets/scenes for both reconstructions to allow for reproducibility.

Improvements for AI systems

Here are the specific improvements for AI systems based on this research, categorized by the technical aspect of improvement:


)Object-Level Fidelity Enhancement (Geometry & Physics):

  1. Single-Image-to-Metric Reconstruction: Implement a pipeline that reconstructs object geometry from a single photograph and explicitly estimates metric scale using an external calibration board (like ChArUco).

  2. Authoritative Physics Modeling: Author per-object physics parameters (mass, friction, restitution) rather than relying on engine defaults or vision-language model estimations for these properties. This involves using the reconstructed mesh to inform material classes and then authoring friction tables based on those classes.

  3. Signed Distance Field (SDF) Collision Geometry: Use SDF representations for collision geometry rather than relying solely on convex decompositions, as SDF is superior for accurately modeling concave shapes (like bowls and mugs), leading to more physically accurate interaction models in simulation.

  4. Scene Fidelity Enhancement (Environment Reconstruction):

  5. Authoritative Scene Splatting: Implement a custom scene reconstruction pipeline that generates the 3D scene splat using the same inputs (handheld video) as the object reconstructions, ensuring temporal and spatial consistency with the authored assets.

  6. Metric-Aligned Environment: Ensure the environment reconstruction is precisely aligned (position and scale) to known real-world measurements (e.g., table plane measurement), rather than relying on unscaled generative models.

  7. Holistic System Improvement (Reducing Sim-to-Real Gap):

  8. Unified Reconstruction Pipeline: Integrate object reconstruction, scene reconstruction, metric scaling, and authored physics into a single authored pipeline to ensure that improvements in one area (e.g., geometry) are not masked by defaults in another (e.g., physics or scene).


Improved AI System Capabilities:

The resulting improved AI system (robot policy running in the simulation) will exhibit the following capabilities:

  1. Enhanced Robustness to Real-World Discrepancies: The robot policy will maintain significantly higher performance fidelity when deployed in a real-world environment compared to policies trained or tested against default simulation assets.

  2. Reduced Score Error and Progress Disagreement: The system's simulated performance scores will be closer to the actual real-world scores (reducing mean score error by up to 10.56 percentage points) and its progression through complex tasks will more accurately reflect real-world capabilities (reducing progress disagreement by up to 8.32 percentage points).

  3. Improved Task Success Rate: The policy is expected to successfully complete a wider variety of manipulation tasks, especially those involving concave objects (bowls, mugs), where the default system struggled due to inaccurate collision modeling and object representation.

  4. Better Policy Ranking Preservation: The simulation will better preserve the relative performance ordering of different robot policies (as evidenced by higher correlation scores, r=0.90 vs r=0.51 for the default), allowing for more reliable policy selection and comparison during development phases where real-world testing is impractical.

  5. Mitigation of Control Gaps: By ensuring the reconstruction pipeline is consistent across both reconstructions, the system addresses discrepancies arising from visual rendering gaps, object physics mismatches, and scene representation errors simultaneously, leading to a more faithful digital twin for policy evaluation.

Abstract

Simulated evaluation is increasingly used alongside real-world evaluation of robot policies because it is cheaper and easier to repeat; however, its value depends on how closely its outcomes track the real robot's. We test whether our reconstruction pipeline, combining metrically scaled object geometry, authored physical parameters and scene reconstruction, reduces disagreement between simulated and real robot scores relative to a default open-source recipe. We constructed two simulated versions of one bimanual robot cell: an authored reconstruction, using object geometry at estimated metric scale, projected textures, authored physics and our own scene splat; and a baseline, referred to as the default reconstruction, using the open-source recipe of a generative single-image mesh, engine-default physics and a Gaussian-splat scene. Both reconstructions use the same object photographs and scene video. Two policies ran five tasks each, giving ten task-policy pairs, which we call cells; each cell was run twenty times in each reconstruction with all other settings held fixed. Both reconstructions were scored against the same real trials, graded by a third-party evaluator. Pearson correlation between the ten simulated and real cell means is r = 0.90 for the authored reconstruction and 0.51 for the default. Mean score error is 6.97 percentage points for the authored reconstruction and 17.54 for the default, a reduction of 10.56 percentage points. These results show that improving the quality of the environment reconstruction through higher visual fidelity, authored physics and metric scale makes the simulation more faithful to the real world and narrows the sim-to-real gap. We release the harness, the per-trial scores, every reported run's configuration, and the assets and scenes of both reconstructions.

Sources

Related papers