CoDimRecon: Agentic Reconstruction of Sim-Ready 3D Scenes with Deformable Curves, Surfaces, and Volumes
cs.CV, cs.GR, cs.RO
Submitted: 2026-09-28
Updated: 2026-10-03
Code: https://github.com/LiteReality/LiteReality-Agent
Project page: https://shuzhaoxie.github.io/CoDimRecon
Terminology
Sources
- SAM 3: Segment Anything with Concepts
- Engine-Native Editable 3D World Reconstruction with Objects and Lighting
- ReplicateAnyScene: Zero-Shot Video-to-3D Composition via Textual-Visual-Spatial Alignment
- LiteReality: Graphics-Ready 3D Scene Reconstruction from RGB-D Scans
- SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image
- $\phi$-Scene: Physically Grounded Image-to-3D Scene Reconstruction
- Advances in 3D Generation: A Survey
- Gaussian Object Carver: Object-Compositional Gaussian Splatting with surfaces completion
- REST3D: Reconstructing Physically Stable 3D Scenes from a Single Image
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- PhyRecon: Physically Plausible Neural Scene Reconstruction
- Decompositional Neural Scene Reconstruction with Generative Diffusion Prior
- SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes
- Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling
- SAM 3D: 3Dfy Anything in Images
- The Replica Dataset: A Digital Replica of Indoor Spaces
- BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender
- VGGT-$\Omega$
- TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation
- FIRE3D: Feed-forward Interactive 3D Scene Reconstruction Within A Minute
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models