Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
cs.CV, cs.CL
Submitted: 2026-09-29
Updated: 2026-09-29
Project page: https://cvlab-kaist.github.io/Imagine3D-LLM
Terminology
Sources
- C3G: Learning Compact 3D Representations with 2K Gaussians
- Qwen2.5-VL Technical Report
- Grounded 3D-LLM with Referent Tokens
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
- Unsupervised Semantic Segmentation by Distilling Feature Correspondences
- Gaussian Error Linear Units (GELUs)
- PF3plat: Pose-Free Feed-Forward 3D Gaussian Splatting
- G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- 3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
- GPT-4o System Card
- GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens
- AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views
- LLaVA-OneVision: Easy Visual Task Transfer
- Language-driven Semantic Segmentation
- Do 3D Large Language Models Really Understand 3D Spatial Relationships?
- SQA3D: Situated Question Answering in 3D Scenes
- SpatialLM: Training Large Language Models for Structured Indoor Modeling
- GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models