GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation
cs.CV, cs.CL
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: Accepted to EMNLP 2026 Findings
Project page: https://geoagent-benchmark.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Modern Vision-Language Models (VLMs) perform well above the human baseline in image geolocalization, a task critically important in disaster response, OSINT verification, and location privacy.
Terminology
Abstract
Modern Vision-Language Models (VLMs) perform well above the human baseline in image geolocalization, a task critically important in disaster response, OSINT verification, and location privacy. However, most efforts to study AI behavior on the task remain limited to static image-based retrieval, classification, and predictions. We argue that faithful recreation of the task should involve embodied navigation, where a multimodal agent autonomously explores its surroundings to gather observations before submitting a prediction. To this end, we introduce GeoAgent, an agentic environment-based benchmark that requires agents to navigate Street View environments to refine their geolocalization through sequential reasoning. Our analysis shows that modern VLMs struggle to discern regional patterns while succeeding at country- and continent-level predictions. When compared to static image-based baselines, agentic navigation significantly improves accuracy across established metrics. We also note severe bias in a developed/developing region context across frontier model architectures and poor self-improvement capabilities given incorrect priors. Overall, our work establishes the challenges of embodied navigation and geospatial reasoning. We publicly release our code and the GeoAgent environment: https://geoagent-benchmark.github.io
Sources
- GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
- Mind2Web: Towards a Generalist Agent for the Web
- A survey of agent interoperability protocols: Model Context Protocol (MCP), Agent Communication Protocol (ACP), Agent-to-Agent Protocol (A2A), and Agent Network Protocol (ANP)
- PIGEON: Predicting Image Geolocations
- Evaluating Precise Geolocation Inference Capabilities of Vision Language Models
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language Models
- From Pixels to Places: A Systematic Benchmark for Evaluating Image Geolocalization Ability in Large Language Models
- Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
- Geolocation with Real Human Gameplay Data: A Large-Scale Dataset and Human-Like Reasoning Framework
- DeepGeo: Photo Localization with Deep Neural Network
- GeoRC: A Benchmark for Geolocation Reasoning Chains
- WebWalker: Benchmarking LLMs in Web Traversal
- Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards
- ReAct: Synergizing Reasoning and Acting in Language Models
- GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models
- NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization
- WebArena: A Realistic Web Environment for Building Autonomous Agents
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models