UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City".
Jane: Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: We’ve been talking through the details of "UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City," and now we need to think about what this means for the future of AI agents. We're looking at how this research frames the relationship between local sensing and overall spatial capability.
Jane: It really boils down to understanding that urban agency isn't something you can just guess from seeing things well in one spot; it depends on whether that initial local evidence stays useful after the agent starts moving through the city.
Meng: If we take this seriously, the practical implication is that any deployment of an AI agent in a city has to account for plan failure and required replanning when execution deviates from initial assumptions. That demands a much smarter planning module than we have right now.
Lu: The authors conclude that agents recognize local evidence more reliably than they orient it, execute useful local motion more reliably than they complete a route, and preserve compliant movement more reliably than they update a plan after the city changes.
Lalam: I think what this tells us for the culture of agent development is that we need to focus heavily on building reliable spatial estimation first, because that seems to be where the most significant failures occur during extended interaction.
Tom: So, in simple terms, we're seeing that current MLLMs are great at recognizing local details but struggle with turning those details into a sustained plan when the environment keeps changing. It’s about maintaining a stable spatial estimate after the initial view vanishes and then constantly revising it as things become infeasible.
Jane: That makes sense; they can see the immediate street but they don't always know how that street connects to the next ten blocks reliably without that sustained spatial grounding. It moves us from simple recognition to actual navigation.
Meng: For engineering, this points toward a need for continuous feedback loops where the agent constantly checks its spatial hypothesis against new sensory input and adjusts its trajectory immediately if the plan becomes obsolete. That’s a lot of processing overhead we have to manage effectively.
Lu: The authors are pointing toward a future where agents must be capable of maintaining that spatial estimate after the current view disappears, and then revising it whenever earlier plans cease to be feasible, which is a much more sophisticated requirement for urban interaction.
Lalam: For applications, this means we need agent frameworks that prioritize persistent spatial awareness over just short-term task completion metrics when operating in complex territories like a big city.
Tom: So, the authors are really saying that city-scale agency requires keeping track of where you are relative to the whole map even when you’re moving, and constantly updating that map as new information comes in. That’s a key distinction we should keep focused on.
Conclusion: Tom: So, we've been diving deep into "UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City," and now we need to look at what this whole thing actually means for us.
Jane: Yeah, it really helps to put the title in context; it’s about figuring out how an AI agent can move beyond just seeing things locally and actually navigate a big city successfully.
Lu: From my side, I think the core idea is testing if that initial visual information holds up when the agent has to actually start moving and dealing with real-world constraints.
Meng: I'm curious about the practical impact—does this mean we can finally deploy agents that handle complex, extended urban tasks without constantly breaking down?
Lalam: I think what this paper really pushes is the idea of building reliable spatial estimates that don't just exist for a second but persist throughout a long journey.
Tom: Exactly! The authors are showing us that success in urban environments isn't about having perfect local vision; it’s about maintaining a good sense of where you are as you explore.
Jane: It seems like they’re moving the focus away from just recognizing things to actually understanding how those things relate to one another across a larger area.
Lu: The five-level ladder they use for evaluation is fascinating because it systematically shows exactly at which point the local perception starts failing when complexity increases.
Meng: That systematic testing is important; it gives us a clear benchmark for what we need to build in our next generation of agents.
Lalam: It suggests that future AI cultural improvements will rely less on instantaneous reaction and more on building robust, persistent mental maps of the world.
Tom: Right, so the main point is that for an agent to be truly useful in a city, it can't just be good at looking around; it has to know how to use what it sees over time.
Jane: It’s a really encouraging look at where this research is headed and what kind of persistent understanding we might achieve with AI agents in the future.
Lu: It opens up so many creative avenues for how AI can perceive and interact with physical reality on a more nuanced level than we currently imagine.
Meng: I'm still thinking about the engineering challenge of building that continuous spatial update mechanism; it sounds like a lot of computational work is involved.
Lalam: That persistent spatial awareness could fundamentally shift how we build interactive systems, making them much more intuitive for users navigating physical spaces.
Tom: So, the authors are showing us that city-scale agency depends on maintaining a spatial estimate after the current view disappears, then revising it whenever earlier plans cease to be feasible.
Jane: That’s a powerful concept because it moves past just solving one problem to building an agent capable of handling continuous change in its environment.
Lu: It really sets a high bar for what we consider 'agency' in autonomous systems operating outside of controlled labs.
Meng: If we can solve that spatial persistence issue, the implications for robotics and even sophisticated user interfaces are huge.
Lalam: These findings suggest that the next big step in AI advancement will be about making those internal world models genuinely persistent and adaptive to dynamic city conditions.
Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang
Shanghai Jiao Tong University National University of Singapore The Chinese University of Hong Kong Shanghai University University of Oxford
cs.CV
Submitted: 2026-08-27
Updated: 2026-09-28
Comments: 36 pages, 11 figures, 8 tables. Project Page: https://urbanground.github.io, Code Repository: https://github.com/UrbanGround/UrbanGround
Code: https://github.com/UrbanGround/UrbanGround
Project page: https://urbanground.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 92/100
The gist: Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move.
Key concepts
- URBANGROUND
- A real-scale simulation environment built from Hong Kong's 3D geospatial data. It allows an AI agent to interact with a city view in a closed loop, testing how local information remains useful after the agent starts moving.
- Spatial Agency Evaluation Ladder
- A five-level test structure used to measure agent performance. Each level increases the complexity of the spatial task, revealing at which point local understanding fails to support sustained goal-directed behavior in a city setting.
- Local Grounding vs. Orientation
- The study distinguishes between an agent's ability to recognize objects locally (local grounding) and its ability to understand where it is in relation to the whole city (orientation). The paper finds local grounding is more reliable than orientation when movement begins.
- Urban Agency
- The capacity of an AI agent to successfully navigate and achieve goals within a complex, real-world urban environment. The research concludes that true agency requires maintaining a spatial estimate and revising plans as the environment changes.
Terminology
Summary
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. This paper investigates how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city by proposing URBANGROUND, a sandbox built from territory-wide 3D geospatial data.
The gist
Current MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable, with their central failure emerging over extended exploration where local abilities do not compose into sustained goal-directed behavior.
URBANGROUND Framework
URBANGROUND is a real-scale environment built from territory-wide 3D geospatial data of Hong Kong, designed to test if local evidence remains useful after an agent begins to move. It supports closed-loop interaction from a first-person view and provides an interactive map for navigation. The framework realizes the interaction loop through three layers:
-
The geospatial layer anchors the visible city and pedestrian connectivity in one geographic frame, recording trajectories in geographic coordinates.
-
The simulation layer implements continuous physical motion with collision, controlling embodied movement, time-of-day systems (day to dusk night), weather systems (rain and fog), and animated pedestrian populations drawn from the Microsoft Rocketbox avatar library.
-
The agent layer connects an external MLLM to the running Unity environment via a client-server interface, exposing model-visible observations plus a structured action interface.
Spatial Agency Evaluation Ladder
The study organizes tasks into a five-level ladder that increases the spatial state required for success while keeping the interaction interface fixed. This ladder is designed to reveal whether evidence grounded at one viewpoint remains useful after the agent begins to move:
-
Level 1 (Local Environment Understanding) probes
Visual recognition,
Orientation understanding,
andActive exploration questions
(RQ1). -
Level 2 (Navigation under Explicit Instructions) tests tasks like
Short-range goal navigation
andLong-range goal navigation
(RQ2). -
Level 3 (Exploration under Implicit Instructions) includes tasks such as
Place-type search
andImplicit intent inference
(RQ2). -
Level 4 (Multi-Task Planning) evaluates tasks like
Time-window scheduling
andMulti-stop route planning.
-
Level 5 (Dynamic Environment Interaction) introduces dynamic changes, including
Dynamic road-closure replanning
andNavigation among pedestrians
(RQ3).
Research Questions and Findings
The analysis follows three research questions to map the growth of the spatial problem:
(RQ1) Can MLLM agents establish a usable local spatial grounding?
Current MLLM agents show reliable performance on short-range atomic spatial reasoning tasks, but Directional grounding is therefore a shared atomic weakness that appears before long-horizon navigation is required.
Furthermore, real-world movement constraints are not maintained as a persistent part of spatial reasoning when they compete with rapid task completion,
as agents often sacrifice road compliance to answer the user more quickly.
(RQ2) How does local grounding scale into goal-directed navigation?
Performance collapses when the route spans only a few city blocks.
Agents cannot compose short-range atomic capabilities into long-range exploration because errors accumulate during execution without effective correction and ultimately cause the task to fail.
Navigation breaks down when agents must convert the user’s request into a stable goal and continue following that goal as the observation changes,
exposing unreliability in instruction understanding.
(RQ3) How do changes in the city affect grounded perception and action?
Changes in visibility, such as dusk and night, weaken the ability of MLLM agents to answer urban questions reliably.
Furthermore, when a route becomes invalid during execution due to road closures or moving pedestrians, agents often continue to produce locally compliant movements even after their spatial plan has become obsolete,
exposing a gap between plausible local behavior and effective adaptation.
Conclusion
The results show that urban agency cannot be inferred from isolated perception or success over a short route.
Current MLLMs possess useful local spatial skills, but these skills lose reliability as interaction extends through the city. City-scale agency depends on maintaining a spatial estimate after the current view disappears, then revising it whenever earlier plans cease to be feasible.
The paper concludes that agents recognize local evidence more reliably than they orient it, execute useful local motion more reliably than they complete a route, and preserve compliant movement more reliably than they update a plan after the city changes.
Experimental Setup
The study evaluates contemporary MLLMs from families including GPT (GPT-5.5, GPT-5.4, GPT-5.2), Claude (Claude-Opus-5, Claude-Opus-4.
Improvements for AI systems
Based on the provided research paper, here are specific improvements that can be made to existing AI systems, along with what these improved systems would be capable of:
-
Acknowledge and integrate
Spatial Agency
as a core competency rather than treating it as an emergent property of visual recognition. -
Implement a closed-loop feedback mechanism that explicitly checks the persistence and reliability of local spatial grounding after every movement, especially when the view changes (e.g., turning corners).
-
Develop robust
state revision
capabilities: When new evidence invalidates an earlier assumption (e.g., a landmark disappearing), the system must actively revise its working spatial state to maintain coherence with the original goal, rather than accumulating errors without correction. -
Enhance pedestrian-aware movement policies: Improve the ability to distinguish between
local compliance
(staying on a path) andsafe progress
(moving toward the goal while avoiding dynamic obstacles), ensuring that local adherence does not lead to catastrophic failures when route availability changes. -
Develop multi-step planning and sequence composition skills: Move beyond short, atomic abilities by enabling agents to compose multiple local spatial actions into sustained, goal-directed behaviors across several city blocks without losing the global objective.
The improved AI system (based on the URBANGROUND framework) would be capable of:
-
Maintain a reliable, coherent map context throughout extended exploration in complex, real-scale urban environments (Hong Kong).
-
Execute sustained navigation tasks across multiple city blocks and complex routes where the path is not immediately visible or explicitly defined.
-
Successfully complete long-range goals by composing short-range spatial reasoning into a valid global trajectory, even when faced with accumulating errors from local perception updates.
-
Adapt reliably to dynamic environmental changes, such as sudden road closures or moving pedestrian traffic, by revising its internal spatial plan in real-time to find a viable alternative route without abandoning the overall objective.
-
Perform complex urban planning tasks, such as multi-stop route optimization (visiting several points in an unsequenced manner), respecting temporal constraints and dynamic changes simultaneously.
Sources
- NaVILA: Legged Robot Vision-Language-Action Model for Navigation
- MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft
- IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
- VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
- GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
- SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds
- UrbanWorld: An Urban World Model for 3D City Generation
- CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
- Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds
- A Survey on Agentic Multimodal Large Language Models
- VideoGameBench: Can Vision-Language Models complete popular video games?
- Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models