LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation

arXiv:2602.02220 · cs.CV, cs.RO · Submitted 2026-02-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation".

Tom: Language-conditioned goal navigation (LGN) requires agents to locate userspecified targets without step-by-step guidance, and this paper introduces HieraNav,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So we've been talking about this new paper on arXiv called "LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation," and what I'm seeing is that it tackles a really tricky area in navigation where agents have to find things just from a language instruction without any step-by-step help.

Jane: Exactly, Tom, and the main thing the paper claims is that existing benchmarks often struggle because they stick to category-level goals or rely on descriptions made by vision models, which can be messy and confusing.

Lu: That's a huge gap in current research; having systematic evaluation across different semantic levels seems crucial for building truly robust goal navigation systems.

Meng: From an engineering standpoint, I'm curious how they managed to create such a comprehensive benchmark without needing complex depth sensors or detailed three dee scene representations.

Lalam: It seems like the core of this work is providing a really solid, real-world testbed for language-conditioned goal navigation.

Tom: Right, and what makes LangMap stand out is that it introduces HieraNav, which lets agents navigate to targets specified at four different semantic levels: scene, room, region, and instance.

Jane: That hierarchical structure is really smart because it allows the agent to understand the target context from broad descriptions down to very specific details.

Lu: I'm fascinated by how they structured those levels—scene for a general object like an 'armchair,' and then drilling down to an instance description like 'square coffee table' or 'armchair beside the bed and balcony.'

Meng: That level of specificity sounds demanding, especially for a system that operates without depth or object masks.

Lalam: The real strength of LangMap is that these annotations are human-verified through a rigorous contrastive protocol where annotators compare things within the same scene to write detailed descriptions covering four hundred fourteen object categories.

Tom: And the results on those annotations are pretty telling; they found that instance descriptions actually outperform GOAT-Bench annotations by twenty-three percentage points in text-to-view matching while using fewer words.

Jane: That comparison really highlights how much more effective and concise human descriptions can be when it comes to identifying specific objects.

Lu: It suggests that the quality of the language grounding is where the real power lies, not just in having a large dataset of images or categories.

Tom: And this leads us nicely into what this means for future systems; we've got this new benchmark, LangMap, and a baseline called PlaNaVid that uses Bounded Diverse Memory to plan its routes based on those rich descriptions.

Jane: PlaNaVid is interesting because it decouples the long-horizon planning from the actual reactive navigation policy using a Qwen2 point 5-VL7B-Instruct planner and Uni-NaVid for movement.

Meng: I'm looking at that architecture, and since it runs without depth or three dee scene representations, how does that memory system handle maintaining context across those navigation steps?

Paper summary: Lalam: The Bounded Diverse Memory component maintains a set of RGB memory states using a "global-uniform update" strategy to keep the temporal coverage consistent under a fixed budget.

Tom: That sounds like an important mechanism for ensuring the agent doesn't lose track of where it is while trying to satisfy multiple goals across those four semantic levels.

Jane: It really shows how they are leveraging context retrieval to prime the reactive policy for navigation, which seems key for complex tasks like multi-goal completion.

Lu: The paper points out that systematic evaluations on LangMap show clear benefits from having both memory and richer context, even when using zero-shot or supervised models.

Tom: But they also flagged some real challenges in their evaluation, specifically mentioning issues with long-tailed categories, small objects, distant targets, and multi-goal completion itself.

Meng: Those limitations are very practical; if an agent struggles with a distant target or a very rare object category in the real world, that's where the system will likely fail.

Lalam: The authors explicitly state that performance tends to be higher when dealing with coarser levels of description, and they noted that finer disambiguation at the instance level presents a significant hurdle.

Jane: So, while it's great for setting a rigorous testbed, it also gives us clear targets for where the next generation of AI needs to improve its ability to handle fine-grained detail.

Tom: Thinking about the bigger picture, if we can reliably benchmark agents on these complex hierarchical goals across real-world data like LangMap, it sets a much clearer path for developing more capable embodied AI.

Lu: The implication is that language-conditioned goal navigation moves from being a theoretical exercise to something that can be systematically tested and improved against concrete standards.

Meng: Practically, this means we can start training agents with these specific challenges in mind, focusing on overcoming those long-tail issues you mentioned.

Lalam: And from a cultural perspective, if these systems become better at understanding hierarchical context through language alone, it could mean more intuitive and less cumbersome interactions with our physical environments.

Jane: It’s exciting because we are moving toward a system that doesn't just recognize an object; it understands where that object is in relation to the room and the scene simultaneously.

Tom: So, to wrap up this discussion on "LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation," we see a powerful combination of detailed human annotation and hierarchical goal specification.

Lu: It really establishes a much higher bar for what we expect from language-conditioned agents in physical spaces.

Meng: For implementation, the challenge remains making that robust planning system work reliably under those real-world constraints without needing extra sensory input like depth maps.

Lalam: It's a solid foundation, proving that human-verified semantic labels can create incredibly rich learning signals for these navigation tasks.

Conclusion: Tom: So we've just been digging into how this paper, "LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation," sets up a really rigorous testbed for agents navigating the real world using language instructions.

Jane: Exactly, and it seems the whole point of this work is to give us a way to systematically evaluate if an agent can actually figure out complex goals just from natural language without having any step-by-step instructions.

Lu: I think what’s really compelling here is how they broke down the navigation task into those four semantic levels—scene, room, region, and instance—which opens up so many creative avenues for how agents could process spatial information.

Meng: From an engineering standpoint, I'm focusing on the implication of having such a detailed benchmark; it tells us exactly what kind of robust capabilities we need to build into these navigation systems if we want them to work reliably in actual environments.

Lalam: And from my perspective as the language model, this means that if we can train an AI on this kind of rich, human-verified data, it could fundamentally improve how our machines interact with physical spaces by giving them a much deeper understanding of context and relationships.

Tom: That's the big picture, Lalam; it’s not just about moving from A to B anymore; it’s about truly understanding the spatial context behind that move.

Jane: And thinking about those authors, they really focused on making sure their annotations were high-quality through a specific protocol, which is crucial because if the benchmark itself is weak, the results we get won't be trustworthy.

Lu: Their methodology for annotation was pretty rigorous; using that contrastive protocol to get detailed descriptions for four hundred fourteen object categories sounds like a very smart way to capture real-world nuance.

Meng: I wonder how this quality translates into practical application; if an agent can accurately distinguish between an 'armchair in the bedroom' and just 'an armchair in a room,' that’s a huge step for robotics deployment.

Lalam: It suggests that future AI culture will be defined by systems that can handle these subtle, hierarchical distinctions effortlessly, making them much more intuitive companions or tools.

Tom: So we’ve seen the technical setup and the results; now we need to look at what this actually means for how we build these next-generation navigation AIs.

AIML, Adelaide University

cs.CV, cs.RO

Submitted: 2026-02-02

Updated: 2026-10-01

Comments: Accepted to NeurIPS 2026. Benchmark and Code: https://bo-miao.github.io/LangMap

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Language-conditioned goal navigation (LGN) requires agents to locate userspecified targets without step-by-step guidance, and this paper introduces HieraNav, an open-vocabulary LGN task with goals

Key concepts

LangMap Benchmark
LangMap is the first real-world 3D indoor navigation benchmark built from HM3D scans. It includes comprehensive, human-verified semantic labels for over 414 object categories, allowing tasks across all four goal levels to be tested in a realistic setting.
HieraNav Task Specification
HieraNav requires agents to navigate based on four levels of language goals. These range from finding any object category in a scene to locating a unique instance described by detailed contextual relations, testing the agent's ability to understand increasingly specific instructions.
PlaNaVid Architecture
PlaNaVid is an RGB-only baseline for multi-goal navigation. It uses Bounded Diverse Memory (BDM) for long-horizon planning and a reactive policy for navigation, allowing it to handle complex goals without needing 3D scene representations or object masks.

Terminology

Summary

Language-conditioned goal navigation (LGN) requires agents to locate userspecified targets without step-by-step guidance, and this paper introduces HieraNav, an open-vocabulary LGN task with goals specified at four hierarchical semantic levels: scene, room, region, and instance. This work addresses the limitations of existing benchmarks by presenting LangMap, a real-world 3D indoor navigation benchmark with human-verified semantic annotations that supports tasks across all four goal levels.

LangMap Benchmark and Annotation Quality

LangMap is the first real-world 3D indoor navigation benchmark built on HM3D scans, enriched with comprehensive, human-verified semantic annotations and navigation tasks across all four semantic levels. It provides region labels and discriminative region and instance descriptions covering 414 object categories, with 349 categories used to construct over 18K tasks. The annotations are produced through a rigorous contrastive protocol where annotators compare same-category regions and instances within each scene to write concise and detailed descriptions. Quantitative analyses validate this quality, showing that instance descriptions outperform GOAT-Bench annotations by 23 percentage points in text-to-view matching while using fewer words.

HieraNav Task Specification

HieraNav requires agents to navigate to language-specified targets across four levels:

  1. Scene-level: any object of the target category in the scene (e.g., 'armchair').

  2. Room-level: an object of the target category in a specified room type (e.g., 'armchair in the bedroom').

  3. Region-level: an object of the target category in a specific room instance, distinguished from same-type rooms by contextual cues (e.g., 'armchair in the bedroom with a geometric rug').

  4. Instance-level: a unique object instance identified by discriminative attributes or contextual relations (e.g., 'square coffee table', 'armchair beside the bed and balcony').

The task generation process iterates over all object categories instead of random sampling to prevent duplication, yielding about 15K single-goal tasks and 720 multi-goal episodes.

PlaNaVid Baseline Architecture

PlaNaVid is a strong RGB-only baseline designed for multi-goal navigation, combining Bounded Diverse Memory (BDM) with high-level planning to prime a reactive policy for navigation. The framework is decoupled into two stages: memory-guided planning and policy navigation with online memory update. In the first stage, BDM retrieves semantically diverse context to select an initial waypoint and heading for each goal. In the second stage, a reactive navigation policy maps RGB observations to actions, updating M online for subsequent goals. The system utilizes Qwen2.5-VL7B-Instruct [46] as the planner and Uni-NaVid [51] as the navigation policy, operating without depth, 3D scene representations, or object masks.

Bounded Diverse Memory (BDM)

The BDM component maintains a bounded set of RGB memory states to provide long-horizon reasoning context. It performs a global-uniform update to ensure near-uniform temporal coverage under a fixed budget Nmax by executing gap-aligned pruning when the buffer size exceeds Nmax. To maximize semantic coverage, it uses a semantic-diverse retrieval strategy, anchoring the most recent state and iteratively selecting remaining states using a greedy max–min strategy based on cosine similarity of embeddings to approximate max–min dispersion.

Performance and Challenges

Systematic evaluations reveal the benefits of memory and richer context. PlaNaVid achieves top-tier success rates, with its multi-goal success rate (SR) reaching 42.6% and its sequence success rate at k=2 (SeqSR@2) reaching 14.3%. However, analyses identify key challenges for future work, including long-tailed categories, small objects, distant targets, and multi-goal completion. Furthermore, performance generally is higher at coarser levels than at the instance level, where finer disambiguation is required. The paper concludes that HieraNav and LangMap establish a rigorous testbed for language-driven embodied navigation.

How it works

  1. Scene-level: Find the object category.

  2. Room-level: Find the object category in a specified room type.

  3. Region-level: Find the object category in a specific room instance with a discriminative region description.

  4. Instance-level: Find the unique object instance by its discriminative description (e.g., armchair beside the bed and balcony).

The navigation policy is driven by an LLM Planner that selects waypoints and headings based on retrieved, semantically diverse memory states from BDM, which primes a reactive policy for navigation. This decoupled design separates long-horizon planning from reactive navigation.

Improvements for AI systems

Here are specific improvements to AI systems based on the findings in this paper, focusing on leveraging HieraNav, LangMap, and PlaNaVid:


  1. Enhance Navigation Robustness via Hierarchical Goal Specification:

  2. Improve Semantic Grounding and Disambiguation using Human-Verified Annotations (LangMap):

  3. Develop Robust Multi-Goal Planning with Contextual Memory (PlaNaVid):

  4. Implement Open-Vocabulary, Zero-Shot Navigation for Unseen Environments:

Specific Improvements and Capabilities:

  1. Improve Navigation Robustness via Hierarchical Goal Specification:

The system can navigate effectively by interpreting complex, nested natural language instructions across four semantic levels—scene, room, region, and instance. This allows the agent to handle highly contextual requests (e.g., Find the phone on the Bluey bed) rather than just simple object category searches (Find a phone).

  1. Improve Semantic Grounding and Disambiguation using Human-Verified Annotations (LangMap):

The system will generate much more accurate and discriminative descriptions for targets. By utilizing LangMap's human-verified instance descriptions, the agent can distinguish between visually similar objects that VLM-generated descriptions often confuse (e.g., correctly identifying the blue chair vs. the red chair based on fine semantic cues). This drastically reduces semantic errors and ambiguity in goal interpretation.

  1. Develop Robust Multi-Goal Planning with Contextual Memory (PlaNaVid):

The agent will be equipped with a Bounded Diverse Memory (BDM) system that stores a compact, semantically diverse history of past states. When facing a multi-goal episode, the LLM planner uses this memory to select the most relevant preceding context for waypoint and heading selection. This enables long-horizon reasoning and sequential task completion without requiring full trajectory storage, leading to higher success rates in complex sequences.

  1. Implement Open-Vocabulary, Zero-Shot Navigation for Unseen Environments:

The system can be deployed effectively in novel 3D indoor environments (unseen scenes) because the benchmark (LangMap) covers a large number of object categories and real-world scenes, and the method (PlaNaVid) operates on RGB observations only, without needing depth sensors or explicit 3D scene representations. This makes the system highly generalizable for open-vocabulary navigation tasks where prior training data is limited.

Abstract

Language-conditioned goal navigation (LGN) requires embodied agents to locate user-specified targets without step-by-step guidance. However, existing benchmarks largely focus on category-level goals or rely on instance descriptions generated by vision-language models, which often contain ambiguities and semantic errors, limiting systematic and reliable evaluation. We introduce HieraNav, an open-vocabulary LGN task with goals at four hierarchical semantic levels: scene, room, region, and instance. We present Language as a Map (LangMap), the first LGN benchmark to enrich real-world indoor 3D scans with human-verified semantic annotations supporting tasks across all four goal levels. Built on HM3D using a contrastive annotation protocol that compares same-scene regions and instances, LangMap provides region labels and discriminative region and instance descriptions covering 414 object categories and contains over 18K tasks. Each target has concise and detailed descriptions, enabling evaluation across instruction styles. Automated and human evaluations validate our annotation quality: our descriptions improve text-to-view matching accuracy over GOAT-Bench's by 23 points on all shared annotated instances, and an independent human audit yields 92.5% unique-and-correct matches. We also propose PlaNaVid, an RGB-only baseline that combines Bounded Diverse Memory with high-level planning to prime a reactive policy for multi-goal navigation, achieving top-tier success rates without depth, 3D scene representations, or object masks. Further analyses reveal that exploration and hierarchical disambiguation failures become more prominent at finer goal levels, while long-tail categories, small objects, distant targets, timely stopping, and multi-goal completion remain challenging. Benchmark and code: https://bo-miao.github.io/LangMap

Sources

Related papers