Mapping the City Through the Lens of Language Models
Wanqi Liu, Rong Zhao, Zhizhou Sha, Qinyu Cui, Yecheng Zhang
cs.CL
Submitted: 2026-08-07
Updated: 2026-08-11
License: http://creativecommons.org/licenses/by/4.0/
The gist: Language models often complete an underspecified reference to a city with unstated assumptions about urban size, form, infrastructure, environment, and function.
Terminology
Abstract
Language models often complete an underspecified reference to a city with unstated assumptions about urban size, form, infrastructure, environment, and function. We measure those assumptions without naming places. Ten open-weight checkpoints rate anonymized profiles derived from real morphological urban centres across 40 audited indicators and seven domains. The design combines constrained probability-based ratings, prespecified reliability screens, lineage-aware aggregation, multiple population weightings, an independent replication sample, and whole-profile validation. The clearest shared tendency favours urban profiles with larger developed area, faster recent growth, greater mapped infrastructure and non-residential capacity, and less sparse form. Most eligible directions recur in the replication data, and direct ratings of complete profiles show moderate agreement with the indicator-wise construction. Geographic differences shrink after accounting for city scale and development, while reliably measured paired tasks indicate that typicality and desirability are often closely aligned. The framework makes an otherwise vague notion of what models regard as an ordinary city empirically traceable. The resulting evidence delineates a shared yet model-dependent portrait of the city through the lens of language models.
Sources
- The Llama 3 Herd of Models
- CityBench: Evaluating the Capabilities of Large Language Models for Urban Tasks
- Gemma 2: Improving Open Language Models at a Practical Size
- Large Language Models are Geographically Biased
- 2 OLMo 2 Furious
- Qwen3 Technical Report
- GenAI Models Capture Urban Science but Oversimplify Complexity
- Culturally uneven urban perception in large language models
- UrbanPlanBench: A Comprehensive Urban Planning Benchmark for Evaluating Large Language Models
- GlobalBuildingAtlas: An Open Global and Complete Dataset of Building Polygons, Heights and LoD1 3D Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering