Perceive to Hypothesize, Verify to Ground: An Agentic Reasoning Framework for Open-World Geo-Localization
cs.AI, cs.MM
Submitted: 2026-08-30
Updated: 2026-08-30
Code: https://github.com/Arandinglv/GeoPAVE
Terminology
Sources
- MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual Reasoning
- GAEA: A Geolocation Aware Conversational Assistant
- GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
- GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces
- GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
- Geospatial Representation Learning: A Survey from Deep Learning to The LLM Era
- Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
- Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization
- SpotAgent: Grounding Visual Geo-localization in Large Vision-Language Models through Agentic Reasoning
- GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language Models
- VisTIRA: Closing the Image-Text Modality Gap in Visual Math Reasoning via Structured Tool Integration
- GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language Model
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language Models
- Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory
- Reasoning or Pattern Matching? Probing Large Vision-Language Models with Visual Puzzles
- WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- GeoRC: A Benchmark for Geolocation Reasoning Chains
- Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection