Toward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language Models
cs.CV, cs.AI
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
- ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
- SAM 3D: 3Dfy Anything in Images
- Grounded 3D-LLM with Referent Tokens
- Reasoning in Space via Grounding in the World
- MapAnything: Universal Feed-Forward Metric 3D Reconstruction
- 3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding
- SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
- SQA3D: Situated Question Answering in 3D Scenes
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
- SAM 2: Segment Anything in Images and Videos
- VGR: Visual Grounded Reasoning
- Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
- Orient Anything V2: Unifying Orientation and Rotation Understanding
- PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes
- GenPose: Generative Category-level Object Pose Estimation via Diffusion Models
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models