Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation
cs.RO, cs.AI
Submitted: 2025-02-20
Updated: 2026-09-16
Comments: 8 pages, 4 figures
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made them powerful tools in embodied navigation, enabling agents to leverage commonsense and spatial
Terminology
Abstract
Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made them powerful tools in embodied navigation, enabling agents to leverage commonsense and spatial reasoning for efficient exploration in unfamiliar environments. Existing LLM-based approaches convert global memory, such as semantic or topological maps, into language descriptions to guide navigation. While this improves efficiency and reduces redundant exploration, the loss of geometric information in language-based representations hinders spatial reasoning, especially in intricate environments. To address this, VLM-based approaches directly process ego-centric visual inputs to select optimal directions for exploration. However, relying solely on a first-person perspective makes navigation a partially observed decision-making problem, leading to suboptimal decisions in complex environments. In this paper, we present a novel vision-language model (VLM)-based navigation framework that addresses these challenges by adaptively retrieving task-relevant cues from a global memory module and integrating them with the agent's egocentric observations. By dynamically aligning global contextual information with local perception, our approach enhances spatial reasoning and decision-making in long-horizon tasks. The proposed method surpasses previous state-of-the-art approaches by a significant margin on both the HSSD and HM3D benchmarks and demonstrates strong performance on a real robot.
Sources
- On Evaluation of Embodied Navigation Agents
- ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects
- End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
- The Llama 3 Herd of Models
- OpenFMNav: Towards Open-Set Zero-Shot Object Navigation via Vision-Language Foundation Models
- Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
- A Survey on Hallucination in Large Vision-Language Models
- Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
- InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment
- PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
- Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots
- CoNVOI: Context-aware Navigation using Vision Language Models in Outdoor and Indoor Environments
- OpenIN: Open-Vocabulary Instance-Oriented Navigation in Dynamic Domestic Environments
- VoroNav: Voronoi-based Zero-shot Object Navigation with Large Language Model
- Large Language Models for Robotics: A Survey
- Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving