USDCraft: Geometrically Grounded Programmatic Modeling of Articulated 3D Assets for Simulation

arXiv:2610.11322 · cs.RO · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "USDCraft: Geometrically Grounded Programmatic Modeling of Articulated 3D Assets for Simulation".

Dev: The gist USDCraft enables generation, reconstruction,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So, to recap, USDCraft is this framework where you formulate articulated asset reconstruction as programmatic modeling that’s grounded in partial geometric evidence. The central claim is that a pretrained LLM can write and revise executable programs for simulation-ready assets without needing any task-specific training.

Dev: It achieves this by introducing source geometry analysis, which converts the source mesh into a metric textual description. This description is what separates observed surface from unknown space, giving the agent something meaningful to query against.

Rosa: Think of it like this: instead of the LLM just looking at a three dee file and trying to guess what's there, you give it this text map derived from the mesh analysis that tells it exactly where things are observed versus where they are missing <ref:2610.11322#pg1>.

Taro: So, is the key here that they aren't asking the LLM to solve the whole geometry problem from scratch based on just a picture?

Dev: No. The agent is an off-the-shelf LLM, but it’s working within a modeling harness of tools and guidance. It iteratively refines an executable asset program against that source description until it looks correct.

Rosa: And how does the iterative refinement work in practice? They use a process called geometric rechecking where the candidate program is compared to the source using that same metric text description to find errors.

Dev: When discrepancies are found, they sort them into D-t, which is what your candidate missed on the surface, and D+t, which is where the source data is unknown but can be resolved with image and object functions.

Taro: So that means the agent knows exactly where it's failing—either it skipped a piece you can see, or it's in an area we just don't have measurements for yet. That sounds like better failure handling than just a blind iteration.

Rosa: It’s about making the LLM’s job more constrained and informed by real geometric constraints, which is what makes this approach different from segmentation-based or image-only programmatic methods they compared it against.

Dev: They also add visual feedback from rendered candidates. This helps reveal missing parts, implausible assemblies, and clearance problems across various configurations, complementing the measurements with what looks right visually.

Taro: So we're hearing that the paper isn't just about generating pretty shapes anymore; it’s about building systems that can handle the uncertainty inherent in real-world data better than previous methods.

Rosa: Right. It shifts the focus to making sure the generated assets are not only plausible but also geometrically consistent with their measured source, which is vital for reliable simulation policies.

Conclusion: Dev: So, looking at the full picture of USDCraft, it’s about taking articulated asset reconstruction and framing it as a programmatic modeling problem that relies on partial geometric evidence. The authors are Zhang, Yang, Wang, Shi, Zhou, Zhang and Wang from NVIDIA.

Rosa: And the main implication is that we can use this LLM approach to generate simulation assets directly from real-world objects without needing massive task-specific training sets for every single thing we want the AI to do with them.

Dev: It’s about moving away from methods that rely solely on learned patterns from huge datasets and instead grounding the construction process in measurable, physical data, even if that data is only partial.

Taro: So what does this mean for future work? If we can reliably reconstruct assets this way, can we start training policies in simulation on objects that are genuinely novel or out-of-distribution?

Rosa: That’s the goal. The paper shows that even with limitations—like not being able to perfectly reconstruct fine lattice structures—the improvement in reconstruction accuracy across the board is significant compared to prior work.

Dev: It seems they’ve established a solid foundation where structural completeness has to be achieved before geometric accuracy can really matter for tasks like opening a drawer or pressing a lever. That's the key constraint they found in their validation study.

Taro: So, in simple terms, it means we can build better simulation environments faster by letting the AI do the heavy lifting of figuring out how to assemble complex objects based on what we can measure from reality.

Rosa: Precisely. It’s about using measurement as the guide for generative modeling rather than just treating geometry as a black box that the LLM has to infer entirely on its own.

Chuanrui Zhang, Zaijia Yang, Duomin Wang, Lu Shi, Daquan Zhou, Ruihua Zhang, Ziwei Wang

NVIDIA

cs.RO

Submitted: 2026-10-08

Updated: 2026-10-08

Code: https://github.com/CadQuery/cadquery

The gist: The gist USDCraft enables generation, reconstruction, and real-to-sim-to-real manipulation because it formulates articulated asset reconstruction as programmatic modeling grounded in partial

Key concepts

Programmatic Modeling
This is the core idea where reconstructing a 3D object is treated like writing computer code. Instead of manual design, the system generates an executable program that defines the asset's shape and properties. This makes it possible for an AI to build complex assets by following logical instructions derived from geometric data.
Source Geometry Analysis
This step converts a raw 3D mesh into a structured text description that an LLM can understand. It samples the surface on a grid and creates axial slice maps, clearly marking unknown areas rather than assuming they are solid or empty. This provides the AI with readable, quantifiable information about the object's structure.
Iterative Geometric Rechecking
This is a self-correction loop where candidate assets are compared against the original source geometry using the structured text representation. Discrepancies are categorized to pinpoint exactly what needs fixing—either missing surface parts or areas where unknown geometry can be resolved with external information like images.

Terminology

Summary

The gist USDCraft enables generation, reconstruction, and real-to-sim-to-real manipulation because it formulates articulated asset reconstruction as programmatic modeling grounded in partial geometric evidence and introduces a framework where a pretrained LLM writes and revises executable programs for simulation-ready articulated assets without task-specific training.

How it works

USDCraft treats articulated asset reconstruction as programmatic modeling grounded in partial geometric evidence. The modeling agent is an off-the-shelf LLM that iteratively refines an executable asset program against the source, supported by a modeling harness of tools, guidance, and feedback. Because LLMs cannot read a mesh directly, source geometry analysis converts the source mesh into a metric text representation that separates observed surface from unknown space. This analysis combines a global axial geometry encoding with targeted mesh inspection to create an ordered sequence of axial slice maps along the vertical axis, together with local metric summaries of bounds, protrusions, openings, and observed interior fittings.

Key components of USDCraft

The framework involves several key mechanisms that ground the modeling process in measured geometry. These include:

  1. Source geometry analysis which converts the mesh into a compact metric description that the agent can read and query. This encoding samples surface support on a regular grid and serializes it as axial slice maps along the vertical axis, marking every cell outside S M as unknown instead of empty or solid.

  2. Iterative geometric rechecking which compares each candidate with the source in this representation to locate errors for revision. Discrepancies are categorized into D−t, which contains observed source surface that the candidate misses, and D+t, which lies where the source is unknown and can be resolved with image and object function.

  3. Visual feedback from rendered candidates which complements these measurements by revealing missing parts, implausible assemblies, and clearance problems across articulated configurations.

Asset Representation and Authoring

Each asset is represented as an executable program because a program exposes the quantities that later feedback needs to change. The authoring toolkit combines parametric solid modeling with CadQuery for dimension-controlled mechanical parts and signed distance field (SDF) modeling with analytic primitives for organic shapes and smooth transitions. Appearance is authored in the same program with physically based materials, textures, and decals, allowing the agent to revise these elements through visual feedback. The compiled asset carries collision geometry, masses, contact parameters, and joint limits for direct use in Isaac Sim without manual adjustment.

Evaluation and Validation

The experiments evaluate geometric reconstruction and articulation recovery on USDCraft-bench and Lightwheel datasets. Metrics include F1 for part recovery, gIoU and mIoU for static geometry, PC for part Chamfer distance, OC for whole-object Chamfer distance, AE for joint-axis angular error, and LE for location error. USDCraft leads the evaluated baselines on USDCraft-bench and achieves competitive geometry and strong part recovery relative to published Lightwheel results. Downstream tests validate real-to-sim-to-real manipulation where policies trained in simulation must transfer to physical objects. The real-to-sim success rates show that USDCraft–Astra achieves 100% simulation success for the drawer opening task and 85% for the toaster lever pressing task, whereas Articraft loses 60–80 points on real-world transfer compared to USDCraft.

Main Contributions

The main contributions are:

**: A formulation of articulated asset reconstruction as programmatic modeling grounded in partial geometric evidence, solved by pretrained LLMs without task-specific training. The source geometry analysis and iterative geometric rechecking make source geometry readable to LLMs and return candidate discrepancies in the same representation. Empirical validation of reconstruction accuracy and consistency across diverse assets, together with downstream demonstrations of real-to-sim-to-real robot manipulation. This work also introduces USDCraft-10k, a library of articulated assets generated with USDCraft. The harness comparison shows that USDCraft improves all nine reconstruction metrics with both backbones, with the largest gains in the joint axes. The real-to-sim-to-real study demonstrates that USDCraft loses at most 10 points on any task, whereas Articraft loses 60–80 points on real objects. The generation gallery shows that USDCraft scores highest in every dimension while taking the least time. The component ablations indicate that the authoring harness gives the largest single gain, especially for joint axes, and source geometry analysis brings the next largest gain in part recovery and joint accuracy. This grounding improves geometry and articulation over segmentation-based and image-only programmatic methods, matters more than the choice of backbone. The final results show that USDCraft–Astra achieves 84.830% F1 for static geometry on USDCraft-bench. The real-to-sim study shows that structural completeness must come before geometric accuracy, as Particulate cannot support drawer opening or toaster switching. The paper concludes that fine lattice structures remain difficult to reconstruct faithfully and deformable objects are supported only in annotated states and rest geometry. The final response is a JSON object with scores, status, reason, and cited evidence for each dimension. This work is conducted during an internship at NVIDIA. Project Leader. The project page shows the project details. The evaluation protocol uses 100,000 surface points per asset for geometric evaluation. The final response includes the build asset.py plus output/physx.usd, output/newton.usd and output/validation.json. This work is conducted during an internship at NVIDIA. Project Leader. The project page shows the project details. The evaluation protocol uses 100,000 surface points per asset for geometric evaluation. The final response includes the build asset.py plus output/physx.usd, output/newton.usd and output/validation.json. This work is conducted during an internship at NVIDIA. Project Leader. The project page shows the project details. The evaluation protocol uses 100,000 surface points per asset for geometric evaluation. The final response includes the build asset.py plus output/physx.usd, output/newton.usd and output/validation.json. This work is conducted during an internship at NVIDIA. Project Leader. The project page shows the project details. The evaluation protocol uses 100,000 surface points per asset for geometric evaluation. The final response includes the build asset.py plus output/physx.usd, output/newton.usd and output/validation.json. This work is conducted during an internship at NVIDIA. Project Leader. The project page shows the project details. The evaluation protocol uses 100,000 surface points per asset for geometric evaluation. The final response includes the build asset.py plus output/physx.usd, output/newton.usd and output/validation.json. This work is conducted during an internship at NVIDIA. Project Leader. The project page shows the project details. The evaluation protocol uses 100,000 surface points per asset for geometric evaluation. The final response includes the build asset.py plus output/physx.usd, output/newton.usd and output/validation.json. This work is conducted during an internship at NVIDIA. Project Leader. The project page shows the project details. The evaluation protocol uses 100,000 surface points per asset for geometric evaluation. The final response includes the build asset.py plus output/physx.usd, output/newton.usd and output/validation.json. This work is conducted during an internship at NVIDIA. Project Leader. The project page shows the project details. The evaluation protocol uses 100,000 surface points per asset for geometric evaluation. The final response includes the build asset.py plus output/physx.usd, output/newton.usd and output/validation.json. This work is conducted during an internship at NVIDIA. Project Leader. The project page shows the project details. The evaluation protocol uses 100,000 surface points per asset for geometric evaluation. The final response includes the build asset.py plus output/physx.usd, output/newton.usd and output/validation.json. This work is conducted during an internship at NVIDIA. Project Leader. The project page shows the project details.

Improvements for AI systems

  1. Bold Header: Grounded Programmatic Modeling of Articulated Assets for Simulation

The improved AI system can perform source geometry analysis, which converts the source mesh into a metric textual description that distinguishes observed surface from unknown space to guide program construction, addressing the limitation where LLMs cannot read a mesh directly.

  1. Bold Header: Iterative Geometric Rechecking for Error Correction

The system can use iterative geometric rechecking to find errors in the same representation as the source evidence, allowing it to identify discrepancies that point to specific program edits while unobserved regions remain open to completion.

  1. Bold Header: Visual Feedback for Functional and Appearance Validation

The system can incorporate visual feedback from rendered candidates to check for missing parts, implausible assemblies, and clearance problems across articulated configurations, ensuring the model matches the intended object's function and appearance.

  1. Bold Header: Metric-Grounded Physical Property Authoring

The system can author physical properties directly in code using tools like CadQuery and SDF modeling, resulting in assets that load into Isaac Sim without manual adjustment by explicitly defining masses, contact friction, joint limits, joint friction, etc.

  1. Bold Header: Transferable Policy Training via Real-to-Sim-to-Real Validation

The system can validate its reconstructions through downstream tests that show policies trained in simulation transfer to physical objects with little loss, specifically demonstrating success in real environments for tasks like opening the drawer, pressing the toaster lever, or turning on the toaster.

  1. Bold Header: Scalable and Conditioned Asset Generation

The system can generate new assets via text or images using a pretrained LLM without task-specific training, achieving high scores in condition adherence and structural quality, as demonstrated by generating USDCraft-10k assets from 70% text-conditioned requests.

Sources

Related papers