Seeing Is Not Measuring: Tool-Augmented Metric Spatial Reasoning for Vision-Language Models
cs.CV
Submitted: 2026-09-24
Updated: 2026-09-24
Terminology
Sources
- WildDet3D: Scaling Promptable 3D Detection in the Wild
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
- MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
- SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
- ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models
- Toolformer: Language Models Can Teach Themselves to Use Tools
- VGGT: Visual Geometry Grounded Transformer
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
- SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
- Depth Pro: Sharp Monocular Metric Depth in Less Than a Second
- SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL
- Geometrically-Constrained Agent for Spatial Reasoning
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
- Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- Visual Spatial Tuning
- Cambrian-S: Towards Spatial Supersensing in Video
- Open Vocabulary Monocular 3D Object Detection
- ReAct: Synergizing Reasoning and Acting in Language Models
- ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models