Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling
cs.CV, cs.AI
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Project Page: https://lucida-r2s.github.io/
Project page: https://lucida-r2s.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Scan2CAD: Learning CAD Model Alignment in RGB-D Scans
- SceneCAD: Predicting Object Alignments and Layouts in RGB-D Scans
- Qwen3-VL Technical Report
- Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents
- SporeAgent: Reinforced Scene-level Plausibility for Object Pose Refinement
- I Like to Move It: 6D Pose Estimation as an Action Decision Process
- ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language
- Boxer: Robust Lifting of Open-World 2D Bounding Boxes to 3D
- HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- Seed3D 2.0: Advancing High-Fidelity Simulation-Ready 3D Content Generation
- ROCA: Robust CAD Model Retrieval and Alignment from a Single Image
- LRM: Large Reconstruction Model for Single Image to 3D
- 3D-LLM: Injecting the 3D World into Large Language Models
- Cooperative Holistic Scene Understanding: Unifying 3D Object, Layout, and Camera Pose Estimation
- WildDet3D: Scaling Promptable 3D Detection in the Wild
- MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation
- Bottom Up Top Down Detection Transformers for Language Grounding in Images and Point Clouds
- Panoptic Neural Fields: A Semantic Object-Aware Neural Scene Representation
- MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models