AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness".
Dev: Zero-shot vision-and-language navigation in continuous environments (VLN-CE) has recently become feasible with large vision-language models (VLMs), but existing methods suffer from limitations in action space, depth utilization, and memory management.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: Moving onto the title and authors, we’re looking at "AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness." The authors are Yijian Li, Changze Li, Hantian Shi, Jiaying Luo, and Jiyuan Cai.
Dev: I see the focus there is on making zero-shot navigation feasible by changing *how* the VLM interacts with the environment rather than just bigger models alone.
Taro: It seems like they are tackling a fundamental problem: how to give a VLM the ability to reason about continuous space without needing massive amounts of task-specific training data for every single scenario.
Rosa: Exactly, Taro; the authors frame it as rethinking zero-shot VLN-CE by treating navigation as an agentic interface that exposes action, depth, and memory as callable tools. This means we’re not just teaching a policy; we’re giving the model a set of explicit actions it can request from its surroundings.
Dev: That sounds like they're trying to solve the problem of poor action space limitations by letting the VLM directly select a target pixel in RGB observations, which is much broader than choosing from a small set of waypoints.
Taro: I wonder how this direct selection capability plays out when the instruction is vague, and what kind of feedback we get if that selection leads to an impossible or unsafe situation.
Rosa: The paper details the "Action Tool" specifically, which takes a selected pixel and returns either an execution command or feedback after it back-projects the target into a three dee point and performs a geometric safety check to ensure no body-height point falls inside the swept corridor <ref:2606.10577#pg1>.
Dev: That deterministic safety checking is something I’m really interested in; having that hard constraint baked into the tool helps manage some of those failure modes we worry about in control systems.
The paper's summary: Rosa: Now, let’s look at the actual summary of "AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness." They explain that they reformulate zero-shot VLN-CE as a tool-calling interface that exposes depth, memory, and action grounding directly to the VLM.
Dev: It seems like the core idea is to replace learned predictors with these three explicit tools: an action tool, a depth tool, and a selective memory recall tool.
Taro: So the summary highlights that they are addressing three main bottlenecks in existing methods: restricted action spaces from waypoints, ineffective depth utilization because it wasn't explicitly exposed to the VLM for spatial reasoning, and context overload from long histories.
Rosa: That’s right; the paper points out that waypoint predictors restrict the action space by only choosing among a small set of candidates, and they also noted that waypoint-based interfaces take depth inputs as ground truth but don't explicitly expose them to the VLM for spatial reasoning.
Dev: And then there’s memory, which they say is often handled by accumulating long textual or visual histories with substantial irrelevant context, or by retrieving cross-episode experiences, which weakens the zero-shot setting.
Taro: I see how that selective memory tool tries to fix the context accumulation issue by combining a compact trajectory image map with a recall tool that lets the VLM selectively revisit past observations without overwhelming its prompt.
The paper's improvements: Rosa: Focusing on what they actually improved, AgenticNav introduces several specific enhancements. They suggest that replacing the Action Tool with a waypoint predictor reduces performance metrics like SR and SPL, which confirms their point about needing to choose visual targets directly instead of being restricted to learned waypoints.
Dev: I’m paying attention to the depth tool because they show that requesting metric depth at precise image locations allows the VLM to leverage that information more effectively than traditional waypoint predictors or direct depth-image inputs for spatial reasoning.
Taro: The paper also points out the memory mechanism as a major improvement; combining a compact trajectory image map with selective visual recall actually improves long-horizon decision-making while avoiding the accumulation of long historical-frame contexts.
Rosa: And they highlight that by designing this agentic harness, they manage to demonstrate state-of-the-art zero-shot performance on the R2R-CE benchmark under fair VLM backbone comparisons, which is a big win for their methodology.
Dev: It's interesting that the ablation study confirms this; removing the Depth Tool drops SR significantly, showing how crucial that spatial information is for their navigation success.
Conclusion: Rosa: To wrap up our discussion on "AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness," the paper successfully rethinks zero-shot VLN-CE by treating navigation as tool calling, moving away from reliance on trained waypoint predictors.
Dev: It’s clear that this harness removes the dependency on those learned predictors and gives the VLM better access to spatial and contextual information through dedicated tools for action, depth, and memory.
Taro: I think the most important implication is that they show that the design of these action, perception, and memory harnesses is as important as the choice of the VLM itself for foundation-model navigation.
Rosa: Precisely; this paper demonstrates that when you design a robust interface like AgenticNav, it can unlock performance gains even with models like GPT-five point five or Gemini-two point five-Pro, showing strong sim-to-real generalization as well.
Dev: From an engineering standpoint, the deterministic safety checking integrated into the action tool is a huge plus for reliability in continuous environments where things can get messy quickly.
Taro: I just want to reiterate that their findings on the R2R-CE benchmark under fair VLM comparisons really show how much more robust these agentic interfaces are compared to prior methods.
Rosa: That’s right; AgenticNav provides a concrete way forward for getting powerful foundation models into truly continuous and reliable navigation tasks.
Dev: We're looking forward to seeing how this harness integrates into real-world hardware and what kind of latency we can expect when deploying these tool calls in production.
Shanghai Jiao Tong University
cs.RO
Submitted: 2026-06-09
Updated: 2026-10-02
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 87/100
The gist: Zero-shot vision-and-language navigation in continuous environments (VLN-CE) has recently become feasible with large vision-language models (VLMs), but existing methods suffer from limitations in
Key concepts
- AgenticNav
- A lightweight harness that reformulates zero-shot navigation as an agentic tool-calling process. It exposes core navigation capabilities like action, depth, and memory as specific tools that the vision-language model can call upon to interact with the environment.
- Waypoint Predictor Limitation
- Existing methods rely on a learned waypoint predictor, which limits the robot's movement choices to a small set of candidates. This approach fails to utilize metric depth information effectively for spatial reasoning and restricts the VLM's freedom of selection in novel scenes.
- On-demand Pixel-Depth Tool
- This tool allows the VLM to request precise metric depth measurements at specific image locations before taking an action. This enables targeted spatial comparisons, such as verifying if a path is open or checking the reachability of a candidate point, enhancing spatial reasoning.
- Selective Memory-Recall Tool
- This mechanism manages navigation history by creating a compact map image summarizing the trajectory and providing a recall tool for selective revisiting. This prevents overwhelming the VLM's context with irrelevant history while allowing it to retrieve specific past visual observations.
Terminology
Summary
Zero-shot vision-and-language navigation in continuous environments (VLN-CE) has recently become feasible with large vision-language models (VLMs), but existing methods suffer from limitations in action space, depth utilization, and memory management. This paper introduces AgenticNav, a lightweight navigation harness that rethinks zero-shot VLN-CE as an agentic interface between the VLM and the environment by exposing action, depth, and memory as callable tools.
The gist
AgenticNav exposes action, depth, and memory as callable tools to reformulate zero-shot VLN-CE as an agentic tool-calling process.
Limitations of Existing Methods
Existing state-of-the-art methods commonly rely on a learned waypoint predictor, which restricts the action space by only choosing from a small set of candidates. This design also fails to leverage depth inputs effectively, as it does not explicitly expose metric distances to the VLM for spatial reasoning. Furthermore, memory is often handled by accumulating long textual or visual histories with substantial irrelevant context or by retrieving cross-episode experiences, which weakens the zero-shot setting and lacks reliability in novel scenes.
AgenticNav Architecture
AgenticNav provides a lightweight harness that exposes navigation capabilities as tools rather than relying on learned predictors. The paper introduces three primary tools:
-
The first tool is a waypoint-free action tool, which allows the VLM to
directly select a target pixel in RGB observations, converting it into executable motion.
This keeps numerical geometry outside the VLM while preserving its freedom to select any visible location. -
The second tool is an on-demand pixel-depth tool, enabling the VLM to
request metric depth at selected image locations before acting,
allowing for targeted comparisons such as checking if a corridor is open or which candidate point is reachable. -
The third tool is a selective memory-recall tool, which renders navigation history as a
compact map image summarizing the historical trajectory,
paired with a recall tool that allows the VLM toselectively revisit past visual observations without overwhelming the prompt context.
Tool Functionality and Safety
The Action Tool takes a selected pixel and returns either an execution command or feedback. It first back-projects the target into a 3D point, calculates the required heading and distance, and then performs a geometric safety check. The action is deemed safe only if no body-height point falls inside the swept corridor of the robot.
If rejected, it returns Reselect feedback
; if accepted, it executes the continuous action.
Evaluation and Contributions
AgenticNav was evaluated on the R2R-CE benchmark using various VLM backbones, including GPT-5.5 and Gemini-2.5-Pro. The results show that AgenticNav achieves state-of-the-art (SOTA) performance among zero-shot methods, demonstrating improvements over prior methods like SmartWay by 11 percentage points in SR and 13.37 points in SPL when using the same VLM backbone. The main contributions include: (1) presenting AgenticNav as a tool-calling process to remove dependency on trained waypoint predictors; (2) introducing an agentic depth tool that leverages depth information more effectively than waypoint predictors; (3) designing an agentic memory mechanism combining a compact trajectory image map with a selective visual recall tool; and (4) demonstrating SOTA zero-shot performance on the R2R-CE benchmark under fair VLM backbone comparisons.
Ablation Analysis
Ablation studies confirm the necessity of each component. Replacing the Action Tool with a waypoint predictor reduces SR/SPL, confirming that the main policy benefits from choosing visual targets directly instead of being restricted to learned waypoints.
Removing the Depth Tool drops SR significantly, while removing both map image and Recall Tool drops SR further, showing that selective recall adds value beyond a visual map when instructions depend on previously seen signs or landmarks.
The results suggest that the limiting factor in zero-shot VLN-CE is no longer only the reasoning model itself, but also the action, depth, and memory harness through which the model acts.
Conclusion
AgenticNav successfully rethinks zero-shot VLN-CE by treating navigation as tool calling. The design keeps safety checking in deterministic tools while allowing the VLM to decide what evidence to inspect and where to move. These results suggest that the design of the action, perception, and memory harness is as important as the choice of the VLM itself
for foundation-model navigation.
Limitations
AgenticNav remains limited by VLM decision quality; in failure analysis, VLM decision mistakes dominate: they account for 88.9% of simulation failures and 71.4% of real-world failures.
Future work could focus on distilling this agentic tool-use capability into compact models for faster inference.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on the AgenticNav framework, and what these improved systems will be able to do:
)1. Enhanced Generalization in Continuous Environments (Zero-Shot Robustness):
The system will move beyond reliance on trained, domain-specific waypoint predictors.
- This improved system can successfully navigate novel, unseen 3D environments (sim-to-real generalization) and complex continuous navigation tasks without any prior training or explicit waypoint mapping.
)2. Precise Metric Spatial Reasoning via On-Demand Depth Query:
The system will gain the ability to perform fine-grained spatial reasoning by querying metric depth only when necessary.
- This allows the agent to answer critical questions like:
Is this corridor wider than that one?
orIs a doorway closer than an obstacle?
This enables more accurate decision-making regarding reachability and spatial relationships, which is currently difficult for VLMs relying solely on visual cues or fixed waypoints.
)3. Contextual Reasoning with Selective Memory Recall (Avoiding Context Overload):
The system will maintain long-horizon planning capabilities without suffering from context window limitations due to irrelevant historical data.
- Instead of accumulating every frame, the agent can selectively retrieve only the most relevant past visual observations (e.g., a specific sign or landmark) when needed for current decision-making. This ensures that high-level reasoning remains grounded in pertinent information while keeping the prompt context manageable and focused on the immediate task.
)4. Direct Pixel-Level Goal Selection (Bypassing Learned Constraints):
The system will gain the freedom to select any visible, reachable target pixel directly from RGB observations, rather than being restricted to a small set of pre-predicted points.
- This removes the bottleneck of learned action spaces, allowing the VLM to choose a goal location that might be missed by traditional predictors but is clearly visible in the current frame. This leads to more flexible and instruction-following navigation policies.
)5. Deterministic Safety and Execution (Guaranteed Safe Motion):
The system will incorporate a deterministic safety checking mechanism integrated directly into the action execution tool.
- Before any continuous motion is executed, the system can verify geometric safety against known obstacles (like body height). If a proposed action violates safety constraints, it will automatically revert to a
Reselect
state, ensuring that the agent only commits to physically safe and executable motions.
Abstract
Zero-shot vision-and-language navigation in continuous environments (VLN-CE) has recently become feasible with large vision-language models (VLMs). Existing methods typically rely on learned waypoint predictors to propose navigable actions. This limits the model's action space and fails to leverage depth inputs effectively. Moreover, memory is commonly handled by accumulating long textual or visual histories, which overwhelms the prompt context. In this paper, we rethink zero-shot VLN-CE as an agentic interface between the VLM and the environment, and present AgenticNav, a lightweight navigation harness that exposes action, depth, and memory as callable tools. Instead of choosing from predicted waypoints, the action tool allows the VLM to directly select a target pixel in RGB observations, converting it into executable motion. Depth is exposed through an on-demand pixel-depth tool, enabling the VLM to request precise metric distances only where they matter. For memory, AgenticNav uses an agentic memory architecture combining reasoning text history and a compact map image, paired with a recall tool that allows the VLM to selectively revisit past visual observations. On R2R-CE, AgenticNav achieves state-of-the-art (SOTA) zero-shot performance under same VLM comparison, reaching 76% success rate and 66.50% SPL. Ablation experiments validate the effectiveness of our tool and memory designs, while backbone evaluations demonstrate more consistent navigation gains as VLMs improve, highlighting the harness's potential to benefit from future VLM advances. Real-world experiments on two distinct robotic platforms further demonstrate its zero-shot generalization and adaptability.
Sources
- $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models
- MC-GPT: Empowering Vision-and-Language Navigation with Memory Map and Reasoning Chains
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving