GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

arXiv:2608.05948 · cs.AI, cs.CV, cs.RO · Submitted 2026-08-06 · Read on arXiv

Shuai Wang, Yaxin Feng, Xuekun Jiang, Shihan Tian, Ningyu Yan, Xing Shen, Chaoyang Lyu, Hui Wang, Yunsong Zhou, Hanqing Wang, Jiangmiao Pang, Yang Xiang, Xing Gao, Chunhua Shen, Weinan Zhang

Shanghai Artificial Intelligence Laboratory · Hong Kong University of Science and Technology · Shanghai Jiao Tong University · Zhejiang University

cs.AI, cs.CV, cs.RO

Submitted: 2026-08-06

Code: https://github.com/newton-physics/newton

Project page: https://internrobotics.github.io/GAUGE

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 83/100

The gist: The paper introduces GAUGE, a benchmark for evaluating the physical fidelity of both numerical physics engines and generative video world models.

Terminology

Summary

The paper introduces GAUGE, a benchmark for evaluating the physical fidelity of both numerical physics engines and generative video world models. The authors motivate this work by noting that "high visual fidelity does not necessarily imply accurate physical dynamics: a simulator may produce realistic-looking observations while misrepresenting motion, contact, or deformation, thereby encouraging policies to exploit simulation artifacts or yielding incorrect rankings of policy performance. They argue that existing evaluations of physical fidelity are often conducted in isolation and rely heavily on perceptual similarity or human judgments, providing limited insight into which physical principles or parameters are violated."

The paper lists three main contributions:

  • A unified real-world task suite: "We design 22 controlled task families spanning rigid bodies, flexible cables, textiles, and volumetric deformable objects. The suite contains approximately 1,560 motion-capture trials, together with uncertainty estimates and task-specific physical observables."

  • Cross-regime physical parameter annotations: "To the best of our knowledge, GAUGE is the first real-world benchmark dataset to provide task-associated, experimentally characterized parameters spanning rigid-body contact, cloth constitutive response, and volumetric soft-body mechanics within a single standardized collection."

  • Two complementary evaluation protocols: "Using the same real-world experimental foundation, we separately benchmark numerical physics engines and generative video world models. The former measures task-specific sim-to-real discrepancies, while the latter distinguishes agreement with physical-law form from parameter accuracy and temporal stability."

GAUGE contains 22 standardized experimental task families covering rigid bodies, one-dimensional flexible objects, quasi-two-dimensional textiles, and three-dimensional deformable bodies. Specifically, it includes 8 rigid-body tasks, 1 quasi-one-dimensional rope task, 6 quasi-two-dimensional textile tasks, and 7 three-dimensional deformable-body tasks. These tasks evaluate collision, friction, momentum transfer, oscillation, self-contact, curvature, strain, and large deformation.

The benchmark includes multiple material variants per category:

  • Rigid objects: wood, plastic, and metal

  • Textile objects: rayon, satin, uniform cloth, Oxford fabric, synthetic leather, and nylon taslan

  • Three-dimensional deformable objects: both soft and hard variants of foam and rubber

The authors use a millimeter-precision motion capture system to track object trajectories. "For rigid objects, attached markers define a body-fixed frame at the object's geometric center, allowing us to recover its 6-DoF pose over time. For deformable objects, supported by the skeletal tracking functionality, adjacent markers are connected to form a triangulated surface mesh, and their 3D positions are recorded throughout each trial. The system consists of 16 NOKOV Mars9H infrared cameras in a 2 m × 2 m × 2 m motion capture volume operating at 180 Hz with sub-millimeter 3D localization accuracy."

The benchmark provides detailed calibration of the physical properties of all assets involved in each task, including object dimensions, mass, and density:

  • we measure friction and restitution for rigid bodies using inclined-plane and collision tests

  • For textiles, tensile, shear, and bending stiffness are obtained from fabric tests conducted by Style3D

  • For three-dimensional deformable objects, Young's modulus and Poisson's ratio are calibrated using tensile tests and Digital Image Correlation (DIC)

The authors evaluate Isaac Sim (v6.0.0), Genesis (v1.12.0), and Newton (v1.3.0) on 14 representative tasks. Each scene is reconstructed from digital assets and motion-capture initial states. The solvers differ by object class: Rigid bodies use PhysX in Isaac Sim, the native solver in Genesis, and MuJoCo in Newton. Textiles use Surface FEM, PBD, and VBD, respectively, whereas volumetric deformable bodies use FEM, explicit and implicit MPM.

The evaluation uses out-of-the-box performance, meaning physical parameters are set from the GAUGE calibrations, while all other settings remain at their defaults without task-specific fine-tuning.

To compare different object classes, each rollout is represented by a generalized trajectory:

  • Rigid bodies: object position P(t) ∈ R3

  • Textiles: Gaussian curvature K(t) ∈ R Nm on the tracked surface mesh

  • Deformable bodies: areas A(t) ∈ R Nf of triangular faces formed by neighboring markers

The primary metrics are:

  • RMSE: measures frame-aligned error

  • DTW: allows monotonic temporal alignment

Additional physical metrics include:

  • LSD (Longest Stationary Duration) for Newton's cradle: The maximum duration among all time intervals that satisfies the stationary condition, measuring whether the cradle exhibits the expected impulse-like motion

  • MTE (Momentum Transfer Efficiency): The ratio between the momentum of the outgoing ball and that of the incoming ball

  • PD (Period Duration) for the pendulum: The time required for the pendulum to complete one full oscillation

  • EL (Energy Loss): The ratio of the final mechanical energy loss to the initial mechanical energy

The authors evaluate six models on rigid-body scenarios: five image-to-video models (Cosmos3-Nano, Cosmos3-Super-Image2Video, Wan-2.2, Wan-2.7, and Seedance 2.0) and the interactive world model Genie 3. For each task, every model receives the same frontal initial frame and a standardized text prompt describing the scene, object properties, and expected motion. They also conduct a controlled ablation by introducing a negative prompt shared across all models to constrain physically implausible behaviors.

Motion recovery uses SAM3 for segmentation and tracking: the object's 2D pixel coordinates are approximated from the centroid of each frame's binary segment mask. The centroids are converted to metric trajectories using known object dimensions as scale.

The evaluation assesses whether a model has learned physical laws from two complementary perspectives: the structural form of the governing equations and their underlying physical parameters. Metrics include:

  • DE (Dynamic Error): measures consistency with Newton's second law via DE = (1/T) Σ m·ẍ t − f t

  • R2 (Coefficient of Determination): measures how well the expected physical model explains the generated trajectory

  • QFI (Quadratic Form Improvement): measures the additional residual reduction obtained after adding a quadratic term to the expected linear relation, detecting curvature beyond the expected form

  • Isaac Sim gives the lowest errors for slope contact, nonsmooth contact, and turntable motion. Its turntable errors are particularly small, with normalized RMSE and DTW of 0.17 and 0.61, respectively.

  • Genesis performs best on the slope-slider task, reaching 0.58 RMSE and 0.69 DTW.

  • These results show that standard contact and sliding behaviors can be reproduced reasonably well, but the ranking depends on the contact geometry and reference frame.

More demanding events reveal larger gaps:

  • Bouncing ball: Even the best bouncing-ball result is 15.63 times the baseline RMSE and 5.58 times the baseline DTW.

  • Newton's cradle: Isaac Sim and Newton produce zero longest stationary duration, and their normalized momentum-transfer efficiencies are only 0.20 and 0.26; Genesis does not produce a valid rollout.

  • Pendulum: Isaac Sim and Newton obtain normalized periods of 1.10 and 1.09. However, all engines accumulate energy error over the longer rollout, with raw energy-loss values ranging from −0.041 to 0.034.

  • Textile stretching: Isaac Sim and Newton achieve normalized errors close to or below one.

  • Textile bending: the best RMSE and DTW increase to 7.94 and 11.78.

  • Textile flinging: Genesis obtains the lowest RMSE of 8.54, whereas Isaac Sim reaches 128.26.

  • Volumetric deformable bodies: "Genesis gives the lowest errors for stretching, shearing, and twisting, while Newton performs best for bending. Nevertheless, even the best-performing deformable body simulations exhibit errors approximately one order of magnitude above the corresponding real-world baselines."

The overall conclusion: "no engine dominates across rigid, textile, and volumetric regimes. The results instead reveal complementary solver strengths and identify dynamic contact, high-acceleration cloth motion, and three-dimensional deformation as the main sources of residual sim-to-real error."

Tab. 4 and Fig. 4 shows that satisfying the structural form of a physical law and recovering its parameters are distinct capabilities. For slope sliding, different models achieve the lowest QFI for different materials, indicating limited consistency across materials. More importantly, a low QFI does not guarantee a correct acceleration. Specifically:

  • For wood, the best acceleration estimate is 2.06 m/s2 versus the measured 2.58 m/s2.

  • For plastic and metal, however, the closest reported accelerations are only 0.75 and 0.43 m/s2, compared with real values of 2.57 and 2.67 m/s2.

For the bouncing ball: "Cosmos3-Super-I2V with the negative prompt obtains the lowest QFI of 12.50, but its inferred acceleration is only 0.088 m/s2. Seedance 2 gives the closest acceleration among the evaluated models, but its value of 1.84 m/s2 remains far below gravitational acceleration."

  • Six of the ten reported model configurations fail to produce a valid Newton's-cradle sequence, and Wan-2.2 with the negative prompt reaches the highest MTE of 0.76, below the real-world value of approximately one.

  • Pendulum: "Wan-2.2 and Genie3 attain an R2 of 0.99, indicating a strong oscillatory fit, but their periods are 1.93 s and 1.90 s rather than the measured 1.06 s. The closest period is 1.83 s from Wan-2.7, which is still about 73% longer than the baseline."

  • Fig. 4 (c) further shows substantial variation in fitted damping and amplitude, including different damping signs across models.

"The paired settings in Tab. 4 show that the negative prompt does not yield a consistent improvement. For example, it reduces the bouncing-ball QFI of Cosmos3-Super-I2V from 270.69 to 12.50 and enables Wan-2.2 to produce a Newton's-cradle rollout, but it increases the wood slope-slider QFI of Cosmos3-Super-I2V from 13.61 to 569.36. The direction and magnitude of the change depend on both the model and the task."

The authors acknowledge: "First, although the current benchmark covers rigid bodies, textiles, and three-dimensional deformable bodies, its materials and calibrated parameter ranges remain limited. Materials within the same nominal category can differ because of surface treatment, internal structure, manufacturing process, temperature, and wear. They also note that the current world-model track focuses on rigid-body tasks that can be evaluated from two-dimensional image trajectories. This representation is insufficient for textiles and volumetric deformable bodies, whose states involve distributed deformation, depth variation, self-occlusion, and self-contact."

The paper concludes: "Our experiments show that current physics engines have strongly domain-dependent accuracy. Isaac Sim performs well on several rigid contact tasks, Genesis is more competitive on dynamic textile and most deformable body tasks, and Newton is strongest in selected deformation cases; none is consistently accurate across all regimes. Current video world models exhibit a related but different limitation. They can generate trajectories that are visually plausible or well fitted by a physical equation while predicting incorrect accelerations, momentum-transfer efficiencies, oscillation periods, or damping behavior. The authors emphasize that physical fidelity cannot be characterized by visual quality, a single trajectory distance, or parameter estimation alone, and that GAUGE provides a diagnostic basis for improving simulators and world models intended for embodied intelligence."

Improvements for AI systems

Based on GAUGE's findings, here are the specific improvements I can make to AI systems:

  • Energy-conserving contact solvers: Add an energy-tracking term to rigid-body contact resolution so restitution and friction coefficients are corrected per-contact. The improved engine will maintain pendulum energy loss within ±1% over long rollouts instead of accumulating errors of ±0.04.

  • Impact-aware collision cascades: Replace single-impulse contact models with a multi-contact momentum redistribution scheme for chains like Newton's cradle. The improved engine will produce stationary durations matching real 0.1–0.3 s intervals and momentum-transfer efficiencies above 0.90 instead of 0.20–0.26.

  • High-acceleration cloth bending models: Upgrade textile bending stiffness to be strain-rate dependent. The improved simulator will reduce flinging/bending error from RMSE 128.26 to below 10, matching Genesis's performance without sacrificing stretch accuracy.

  • Automated per-regime solver arbitration: Build a router that selects the best solver per object class and task type (Isaac Sim for rigid contact, Genesis for dynamic textiles/deformables, Newton for selected deformations). The combined system will always pick the top-2 engine per task family, eliminating the need for manual solver selection.

  • Physics-constrained trajectory decoding: Add a differentiable physics-loss layer that penalizes violations of Newton's second law, energy conservation, and known pendulum periods. The improved model will predict slopes with acceleration errors under 10% (e.g., 2.3–2.5 m/s2 for wood instead of 2.06) and bouncing-ball accelerations near 9.8 m/s2 instead of 0.088–1.84 m/s2.

  • Explicit physical-parameter heads: Add output heads that predict mass, friction, restitution, and stiffness alongside each generated video. The model will be fine-tuned with the DE, R2, and QFI metrics as losses, enabling it to separate law-form compliance from parameter accuracy.

  • Task-aware prompt conditioning: Replace blanket negative prompts with task-specific conditioning embeddings learned from GAUGE trials. The improved system will avoid regression on slope-slider QFI from 13.61 to 569.36 and produce consistent improvements across all 8 rigid-body tasks.

  • Temporal-dynamics fine-tuning: Fine-tune image-to-video models on motion-capture trajectories to correct period and damping biases. The model will predict pendulum periods within 5% of 1.06 s instead of 1.83–1.93 s, and produce consistent damped oscillation signs across materials.

  • 3D-aware deformable video generation: Extend the world-model evaluation/training pipeline to take multi-view or depth-conditioned inputs so textiles and volumetric deformables are represented as surface meshes or point clouds rather than 2D pixels. The improved model will generate cloth drape and foam compression sequences that match Gaussian-curvature and face-area trajectories within one order of magnitude of real data.

  • Calibration-driven simulator configuration: Use GAUGE's measured material parameters (Young's modulus, Poisson's ratio, friction, restitution, bending stiffness) as priors, then run brief autotuning per task family against the 1,560 motion-capture trials. The resulting simulator will achieve sub-baseline RMSE on 12 of 14 representative tasks without per-task hand-tuning.

  • Cross-regime fidelity scoring: Create a composite fidelity score combining RMSE, DTW, LSD, MTE, PD, EL, DE, R2, and QFI into a single diagnostic profile. The system will automatically identify which physical principle is violated (law form vs. parameter vs. temporal stability) and recommend a targeted fix (solver change, parameter recalibration, or training-data augmentation).

  • Simulation-artifact-aware policy training: Train reinforcement-learning policies with a GAUGE-based reward penalty that detects when the simulator's trajectory deviates from real motion-capture reference (e.g., exploiting incorrect Newton's-cradle momentum). The improved agent will transfer to real hardware without re-tuning, because policies will no longer depend on artifacts like zero-momentum chain reactions or energy-gaining pendulums.


The improved systems can: (1) produce physically accurate rollouts across rigid, textile, and deformable regimes with verified energy and momentum conservation; (2) generate video predictions whose inferred physics parameters match real measurements, not just visually plausible frames; and (3) be used to train embodied policies that transfer reliably to the real world because the simulator's failure modes are known and bounded.

Sources

Related papers