Seek-and-View Reasoning for Multi-View Spatial Understanding
cs.CV
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/q1xiangchen/Vantage
Terminology
Sources
- OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping
- OpenView: Empowering MLLMs with Out-of-view VQA
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- Grounded 3D-Aware Spatial Vision-Language Modeling
- Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views
- Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs
- SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning
- SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
- World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
- Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators
- Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence
- How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
- Think3D: Thinking with Space for Spatial Reasoning
- Depth Anything 3: Recovering the Visual Space from Any Views
- $\pi^3$: Permutation-Equivariant Visual Geometry Learning
- VGGT-$\Omega$
- G3T Up! Gravity Aligned Coordinate Frames Simplify Pointmap Processing
- Gemma 4 Technical Report
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models