HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis
cs.CV
Submitted: 2026-09-22
Updated: 2026-10-05
Code: https://github.com/FishWoWater/CAST
Project page: https://cwchenwang.github.io/harmony
Terminology
Sources
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen2.5-VL Technical Report
- HOLODECK 2.0: Vision-Language-Guided 3D World Generation with Editing
- Segment Anything
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
- Agentic 3D Scene Generation with Spatially Contextualized VLMs
- SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass
- SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes
- SAM 2: Segment Anything in Images and Videos
- SAM 3D: 3Dfy Anything in Images
- 3D-RE-GEN: 3D Reconstruction of Indoor Scenes with a Generative Framework
- ZeroScene: A Zero-Shot Framework for 3D Scene Generation from a Single Image and Controllable Texture Editing
- Qwen3 Technical Report
- Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details
- Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
- Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention
- SAGE: Scalable Agentic 3D Scene Generation for Embodied AI
- Structured 3D Latents for Scalable and Versatile 3D Generation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models