Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation
cs.CV, cs.AI, cs.CL, cs.RO
Submitted: 2026-05-25
Updated: 2026-05-25
Journal ref: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pp. 15243-15252
Code: https://github.com/Teacher-Tom/HSGM_public
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models
- Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- MC-GPT: Empowering Vision-and-Language Navigation with Memory Map and Reasoning Chains
- Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
- NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
- GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
- VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
- OpenAI GPT-5 System Card
- DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
- Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models