You Cannot Photograph the Same Street Twice: Reliability Limits in Vision-Language Measurement of Urban Change

arXiv:2609.00649 · cs.CV · Submitted 2026-09-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "You Cannot Photograph the Same Street Twice".

Tom: Re-photographing an unchanged street moves a perception score by two-thirds as far as changing it,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: The main idea here in "You Cannot Photograph the Same Street Twice: Reliability Limits in Vision-Language Measurement of Urban Change" is that re-photographing an unchanged street moves a perception score by two-thirds as far as changing it, which establishes a really low bar for what we can expect from these models when measuring subtle changes.

Jane: What they claim is that vision-language models are reliable at the scale of hundreds of paired observations but not at the scale of individual sample points, which means relying on just one picture to judge change is risky.

Lu: This finding suggests that aggregating data from many different views and times gives us a much more coherent signal about actual urban development, rather than getting fooled by single snapshots.

Meng: I see how that connects to practical application; if we only look at one street, we might misinterpret natural seasonal changes or minor upkeep as significant redevelopment.

Lalam: Exactly; this paper shows that the utility comes from the volume of data, not just the quality of a single observation, which really shapes how we deploy these AI systems in real-world monitoring.

Conclusion: Tom: When we wrap up our discussion on "You Cannot Photograph the Same Street Twice: Reliability Limits in Vision-Language Measurement of Urban Change," it seems the core message is that reliability in this field depends on aggregation, not just individual data points.

Jane: The authors are showing us that the usable unit for measuring change isn't a single sample point, but rather a few hundred paired observations, which gives us a much more robust picture.

Lu: This has huge implications for how we build systems that monitor cities; it tells us exactly how much data volume is needed to trust the AI's assessment of physical changes.

Meng: For practical deployment, this means engineers shouldn't rely on a single image score when making critical decisions about infrastructure or property condition updates.

Lalam: I think the real impact here is that by understanding these limits, we can design better systems that use these vision-language models in ways that are genuinely useful for city planning and monitoring.

Kaizhen Tan

New York University

cs.CV

Submitted: 2026-09-01

Updated: 2026-10-05

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 83/100

The gist: Re-photographing an unchanged street moves a perception score by two-thirds as far as changing it, establishing that vision-language models are reliable at the scale of hundreds of paired

Key concepts

Perception Score Change
This is the numerical change in how much a vision-language model perceives a street's condition after being shown an image. The paper shows this change is significant when comparing different streets but small when re-photographing the exact same street.
Aggregation
Combining measurements from hundreds of paired observations allows the system to recover a meaningful signal about urban change, such as redevelopment. This aggregation is what makes the overall measurement reliable, rather than relying on any single photograph.
Systematic Drift
A small, consistent error (about 0.1 points) that occurs when comparing measurements of control standpoints over time. This drift grows with the time interval between captures and sets a fundamental limit on how much change the aggregation method can reliably detect.

Terminology

Summary

Re-photographing an unchanged street moves a perception score by two-thirds as far as changing it, establishing that vision-language models are reliable at the scale of hundreds of paired observations but not at the scale of individual sample points.

The Gist

Re-photographing the same street changes a perception score by 0.80 points on average, equivalent to 66.5% of the difference between two different streets in the same city, and aggregation recovers a coherent redevelopment signal across hundreds of paired observations, not individual sample points.

Reliability Limits and Noise Sources

The paper investigates longitudinal reliability by testing how much a perception score can change when the street itself does not undergo substantial redevelopment. The analysis reveals that repeated model calls contribute almost no variation, while image re-encoding and prompt-order changes each account for about one fifth of the between-street difference. Six image statistics—scattering, contrast, colour, exposure, sharpness and specularity—explain almost none of the remaining epochto-epoch variation. A small systematic drift of about 0.1 points remains and increases with the interval between captures, consistent with minor physical changes not recorded by redevelopment labels.

Acquisition Conditions as a Source of Bias

Controlled experiments demonstrate that acquisition conditions can shift scores when camera and image properties are allowed to vary, and that the direction of these shifts depends on the model. In crowdsourced imagery, camera geometry alone causes a model to report physical change in 45% of identical-scene pairs. Normalizing both images to a common virtual camera reduces this rate to 7.5%. Furthermore, "the sign of the acquisition response depends on which model is asked: across three models scored on identical inputs, half the dimensions disagree on sign, and one reverses from significantly negative under one model to significantly positive under another."

The Role of Aggregation and Systematic Drift

Aggregation restores the measurement. Over a few hundred paired observations, standpoints that cross a labelled change point are judged wealthier, better maintained and more enclosed, and less green. The variance measured here converts directly into the number of paired observations an effect of a given size requires. The systematic drift on control standpoints reaches 0.12 points at most, a tenth of the between-street difference, and it grows with the interval between captures rather than with the calendar, setting a floor on what aggregation can recover.

The Usable Unit for Measurement

The usable unit is a few hundred paired observations, not a single sample point. A map in which each sample point carries its own perception change is displaying it at the point level, where a single location’s measured change is dominated by this quantity. The paper establishes a reliability ladder placing sampling noise, encoding noise, prompt-order noise, same-place photographic noise and between-place variation on one scale. Aggregation is the condition under which such a design produces anything.

Model Choice and Directional Correction

The sign of the acquisition response depends on which model is asked. For instance, Beautiful reverses in sign as well, with one of the two models individually significant. This implies that a directional correction cannot be applied without first fixing the model, as a correction estimated with one model may point the other way under another.

Camera Geometry and Normalization

In crowdsourced imagery, camera geometry alone causes a model to report physical change in 45% of identical-scene pairs. However, normalising both images into one virtual camera removes most of the spurious change classification and little of the perception artefact. This normalization reduces the physical change perception by a factor from 0.39 to 0.35, demonstrating that while geometry is a source of artifact, it is not entirely insurmountable when controlled.

Operational Thresholds for Change Detection

The usable unit is defined by the required sample size at which an effect becomes resolvable. For an effect of a tenth of a point, 1,010 labelled-change transitions are required against a control set of the size realized in the study. Below about a tenth of a point, the drift of Section 4.1 becomes the binding constraint rather than the sample size. The operational threshold that separates substantial change from noise is defined by what a planner recording what happened on this street write[s] down.

Limitations and Future Directions

The reliability measurement is made on Google Street View imagery, rendered to a common virtual camera and requested at a fixed heading and field of view, establishing a lower bound for crowdsourced imagery where camera heterogeneity adds the component measured in Section 6. The control standpoints are certified by human labels that mark substantial redevelopment, but the signed drift in Section 4.1 is the size of that contribution, meaning separating it completely from instrument noise would require a labelling scheme with an explicit magnitude threshold applied to the control set.

Improvements for AI systems

Here are specific improvements for AI systems based on the findings of this research, categorized by application:


) Model Calibration and Robustness Improvement:

  1. Improve VLM Reliability with Contextual Filtering: Implement a mechanism that automatically assesses the reliability of a perception score based on acquisition metadata (camera geometry, lighting statistics like contrast/specularity/saturation) and prompt structure.

  2. Implement Model-Specific Sign Correction: Develop a meta-layer that selects the appropriate model configuration (or applies model-specific bias corrections) based on the observed shift in sign from one model to another when scoring identical inputs. This ensures directional corrections are applied consistently across different VLM architectures (e.g., correcting beautiful perception if Llama-4 is positive and Gemini-3.7 is negative).

  3. Integrate Acquisition Noise Decomposition: Create a decomposition module that separates the epoch-to-epoch score variance into components attributable to image statistics (scattering, contrast, etc.), prompt order, and the systematic drift (physical change not recorded by labels). This allows the system to distinguish between genuine urban change signal and measurement noise.

) Urban Change Detection System Improvement:

  1. Shift from Point-Level Mapping to Aggregated Signal Analysis: Instead of reporting a perception score for every single sample point, the system should prioritize aggregating scores across hundreds of paired observations (as suggested by Section 7). This shifts the output from noisy individual points to a coherent redevelopment signal (e.g., identifying an entire street as wealthier or more enclosed).

  2. Establish Reliable Sample Size Thresholds: Implement a dynamic threshold based on the required effect size and the observed systematic drift (Section 7, Table 1). The system should only report change when the magnitude of the score shift exceeds a level determined by this threshold, effectively filtering out noise below the systematic floor.

  3. Incorporate Ground Truth for Physical vs. Perceptual Change: Design a multi-faceted classification system that separates physical change (e.g., new construction) from perceptual change (e.g., lighting shift). The system should be trained to ignore non-substantial changes like repainting kerbs or temporary scaffolding, as defined by the magnitude threshold in Prompt Text C.

) Data Ingestion and Pre-processing Improvement:

  1. Mandatory Image Normalization Pipeline: Before scoring, implement a standardized pipeline that normalizes all input images into a common virtual camera frame (as shown in Figure 8b), explicitly removing geometric variations from different cameras or viewpoints to isolate the perceptual signal.

  2. Metadata-Driven Pair Selection: Utilize advanced tools (like Torkko's approach) to find visually aligned pairs based on refined metadata and feature matching, rather than relying solely on proximity or coarse matching, significantly improving the quality of the longitudinal comparison data used for training/evaluation.

) Operational Output Improvement:

  1. Output Uncertainty Quantification: Every reported change classification should include an explicit uncertainty quantification derived from the acquisition noise decomposition (Section 4.1). This quantifies whether a reported shift is due to real change, model bias, or simply measurement noise below the aggregation threshold.

  2. Develop Model-Agnostic Reporting: The system should be designed so that the final output does not rely on a single VLM's internal sign conventions but instead reports consensus across multiple models (e.g., reporting Wealthy if 5 out of 6 models agree, rather than relying on one model's potentially reversed sign for a specific dimension).

Sources

Related papers