Vision Models Predict Urban Scene Appraisal with Limited Neural Alignment

arXiv:2608.30964 · cs.CV · Submitted 2026-08-31 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Vision Models Predict Urban Scene Appraisal with Limited Neural Alignment".

Jane: Pretrained vision embeddings are increasingly used as general-purpose representations for modelling how people appraise urban scenes, and are validated almost entirely by how well they predict human ratings.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let's talk about the title of "Vision Models Predict Urban Scene Appraisal with Limited Neural Alignment." It immediately tells us the core tension of the research: there’s a prediction task for urban scenes, but there’s a limitation when it comes to matching neural geometry.

Jane: Exactly, Tom. The paper is testing if these powerful vision embeddings truly represent scenes in a way that mirrors how the human brain processes them spatially and temporally, which is what "neural alignment" means in this context.

Lu: What's interesting about the authors Kaizhen Tan and Yuantao Deng is that they are bridging the gap between deep learning representations and neuroscientific data using representational similarity analysis to compare them pairwise.

Meng: That sounds quite complex to set up experimentally; how did they manage to compare those high-dimensional feature spaces with the neural responses without having a perfect shared format?

Lalam: They used representational similarity analysis, which is a way of summarizing both the scenes and the neural responses through their pairwise arrangement, comparing them based on how dissimilar every pair of scenes is according to a set of neural response patterns or model features.

Tom: That’s right; they are essentially measuring how well two different systems—the vision model and the brain—agree on which pairs of images are similar or dissimilar. This sets up the whole experiment for testing that alignment.

Jane: And the implication there is that they are looking at how scene representations unfold over time, showing low-level structure appearing early, and then spatial layout structure emerging by roughly two hundred fifty milliseconds according to EEG data <ref:2608.30964#pg1>.

Lu: That temporal aspect is significant because it suggests the brain builds up understanding layer by layer as we look at a scene, which is something models need to capture in that same hierarchical way.

Meng: So they are mapping out the temporal unfolding of scene understanding versus how fast a model learns those things from its training data.

Lalam: And while they are looking at urban scenes, which differ from standard benchmarks because the targets are evaluative quantities like perceived safety or beauty instead of just object categories, that adds a layer of real-world complexity to the study.

The paper's summary: Tom: Now we move into what the paper actually found in "Vision Models Predict Urban Scene Appraisal with Limited Neural Alignment." Essentially, they tested seventeen different feature spaces across four training objectives to see which one behaved best when compared against neural geometry.

Jane: They looked at everything from language-supervised models like CLIP ViT-B to self-supervised models like DINOv2 ViT-B, and even interpretable controls like Gabor energy descriptors, trying to find a match for human perception.

Lu: The main finding on correspondence is that it’s low throughout the panel; the best representation they tested, DINOv2 ViT-B, only reached twenty-nine point six percent of the lower bound of the noise ceiling when measuring its match to neural geometry <ref:2608.30964#pg0,29.6% of the lower bound of the noise ceiling>.

Meng: So, even the best model isn't perfectly aligned with human brain patterns in this way? That’s a tough result for anyone building models based on these embeddings.

Lalam: It means that even though DINOv2 ViT-B predicts appraisal at up to zero point eight seven, which is quite high, there isn't a reliable relationship between how well it predicts appraisal and how closely it matches the neural geometry across different models.

Tom: That’s the big picture here; prediction accuracy and structural correspondence are independent metrics in this study, which is a really important distinction for anyone trying to validate these tools.

Jane: And they found something really telling about layers: even within a single model family, deeper layers still matched later neural responses, which suggests that "the hierarchical correspondence found for object recognition survives even at this low overall level."

Lu: That is a subtle but important point; it implies that the underlying organization of features, like object recognition structure, does persist in those deeper layers regardless of the overall alignment score.

Meng: From an engineering standpoint, if lower-level structural correspondence is what matters for some tasks, then focusing on those early layers might be more practical than trying to force a perfect match across the entire feature space.

The paper's improvements: Tom: The paper suggests a few ways to think about this research moving forward, and I want to highlight how they propose integrating neural geometry directly into the model design constraints instead of just treating it as an afterthought.

Jane: They argue that we shouldn't rely solely on predictive accuracy for urban scene appraisal; instead, the improved AI system should have a loss function that penalizes deviations from learned representational geometry in the neural response space.

Lu: That means creating multimodal generative architectures where the loss function is augmented not only by standard appraisal prediction error but also by a term that minimizes the distance in the neural RDM space, aiming for perceptually consistent representations.

Meng: That sounds like a very heavy lift for model training; it’s essentially trying to bake brain-like organization into the learning process itself, which is ambitious.

Lalam: This approach could lead to vision models that are not just accurate predictors of ratings but are structurally aligned with how humans organize and perceive scenes in a low-level, brain-like manner, giving us representations where internal distances actually correspond to human perceptual distances.

Tom: And then there’s the calibration point; they propose using the noise ceiling derived from participant agreement as a clear benchmark for measuring correspondence across different models.

Jane: That gives us an objective way to judge quality because we know what level of variability is normal for human perception, allowing us to definitively state whether a new architecture has captured meaningful neural structure or just high-frequency noise.

Lu: By using this noise ceiling as a calibrated denominator, we gain a much more rigorous metric for assessing how good the learned embeddings actually are when compared to biological data.

Conclusion: Tom: So, wrapping up on "Vision Models Predict Urban Scene Appraisal with Limited Neural Alignment," the main message is that we shouldn't assume high prediction accuracy automatically means a model has learned the right way to organize scenes for human perception.

Jane: The paper clearly shows there’s no reliable relationship between how well a model predicts appraisal and how closely it matches neural geometry across different models, which is a crucial distinction we need to keep in mind.

Meng: It sounds like the practical implication is that for applications focused on predicting things like perceived safety or beauty, the held-out predictive accuracy remains the most relevant validation target because structural alignment isn't directly tied to that score.

Lalam: I think for a culture shift, this suggests we need to establish properties of human perception directly through methods like this rather than trying to infer them just from prediction accuracy numbers.

Lu: The work confirms that the neural geometry doesn't depend on the evaluative goal, meaning scene structures are consistent whether participants are judging safety or beauty, which simplifies how we think about scene representation invariance.

Tom: Fantastic points. So, to recap, this paper shows that while models can be highly predictive of appraisal scores, they don't necessarily organize those scenes in a way that mirrors the underlying neural structure we see in brain data from EEG studies on Berlin street scenes.

Jane: It’s a reminder that structural alignment is a separate dimension from prediction accuracy when evaluating vision models for complex tasks like urban scene appraisal.

Meng: If we are building systems for real-world safety applications, focusing on those structural elements, especially at the lower levels where they match neural responses better, could be a more reliable path forward.

Lalam: This research on "Vision Models Predict Urban Scene Appraisal with Limited Neural Alignment" provides a solid foundation for moving toward AI that is not only smart but also structurally informed by how humans perceive the world.

Kaizhen Tan, Yuantao Deng

New York University

cs.CV

Submitted: 2026-08-31

Updated: 2026-10-04

Importance score: 72/100

The gist: Pretrained vision embeddings are increasingly used as general-purpose representations for modelling how people appraise urban scenes, and are validated almost entirely by how well they predict human

Key concepts

Pretrained Vision Embeddings
These are representations learned from large datasets of images without specific task training. They are used as general tools to describe what a scene looks like, and researchers test if these learned descriptions align with how humans perceive urban scenes.
Neural Geometry
This refers to the structure or arrangement of brain activity patterns measured via EEG when people view and rate street scenes. Researchers look for patterns in this brain data to see if they match the way vision models organize visual information.
Representational Similarity Analysis
This is a method used to compare two things—in this case, a scene and a brain response. It summarizes both by looking at how similar their internal arrangements are, allowing researchers to quantify the correspondence between visual data and neural activity over time.

Terminology

Summary

Pretrained vision embeddings are increasingly used as general-purpose representations for modelling how people appraise urban scenes, and are validated almost entirely by how well they predict human ratings. The best representation, DINOv2 ViT-B, reaches 29.6% of the lower bound of the noise ceiling while predicting appraisal at up to 0.87, yet there is no reliable relationship between how well a model predicts appraisal and how closely it matches the neural geometry across models.

How it works

The study tests whether pretrained vision embeddings organize scenes as human perception does by comparing them against representational geometry estimated from brain data (EEG). This comparison is conducted using representational similarity analysis, which involves summarizing both the scenes and the neural responses through their pairwise arrangement. The researchers use openly released EEG from 63 adults who viewed and rated 56 Berlin street scenes to estimate this geometry over time, noting that low-level structure appears early while spatial layout structure emerges by roughly 250 ms.

What was tested

The research addresses three core questions:

  1. How much of the neural representational geometry of urban scenes do current vision representations capture, relative to how much is explainable at all?

  2. How do correspondence patterns vary with model family, scale, training domain and representational depth?

  3. Does better correspondence with the neural geometry go together with better prediction of human appraisal?

The study compares seventeen feature spaces across four training objectives: five language-supervised (e.g., CLIP ViT-B/32), three self-supervised (e.g., DINOv2 ViT-B), three category-supervised (e.g., ResNet-50 trained on ImageNet), three dense prediction models (e.g., SegFormer-B5 fine-tuned on ADE20K), and two interpretable controls (Gabor energy descriptor, colour/luminance statistics).

Key Findings on Correspondence

The overall correspondence between vision representations and neural geometry is low throughout the panel. The best representation, DINOv2 ViT-B, reaches 29.6% of the lower bound of the noise ceiling. A Gabor energy descriptor is indistinguishable from this best model and outperforms every language-supervised model tested. Within a single model family, deeper layers still match later neural responses, suggesting that the hierarchical correspondence found for object recognition survives even at this low overall level.

Key Findings on Prediction vs. Correspondence

A crucial finding is that Strong prediction of human appraisal does not track neural correspondence. The two measures are unrelated across models; the models predicting appraisal best are among the least aligned. Furthermore, attempts to improve alignment by reweighting features towards neural geometry lowered appraisal prediction for every model tested, suggesting that steering features towards the neural geometry overfits rather than recovering a generalisable neural subspace.

Key Findings on Task-Set and Goal Stability

The study demonstrated that the neural geometry does not depend on the evaluative goal. The similarity of scene geometries across different appraisal goals suggests that task-set differences are unlikely to account for the observed model-brain gap. Furthermore, Same-task and cross-task scene geometries were nearly identical once each was corrected for its own reliability, meaning scenes were arranged the same way whether a participant was about to judge safety, beauty, or openness.

Practical Implications

The results suggest that for applications aiming to predict perceived safety or beauty, held-out predictive accuracy remains the relevant validation target. The paper concludes that reading embedding dimensions or embedding distances as properties of human perception should be established directly rather than inferred from prediction accuracy. The benchmark is practical because it requires no training and only 55 images for evaluation.

The gist

The best representation, DINOv2 ViT-B, reaches 29.6% of the lower bound of the noise ceiling while predicting appraisal at up to 0.87, yet there is no reliable relationship between how well a model predicts appraisal and how closely it matches the neural geometry across models.

Table 1: Correspondence between model representational geometry and the neural geometry of urban scenes, in the 150–500 ms window.

Model Layer % ceiling Appraisal r


language-supervised CLIP ViT-B/32 L08 22-0.005 0.80 -

Figure 1: Stimuli, task and the structure of urban appraisal.

(Description of Figure 1: Shows the distribution of nine ratings across scenes, their intercorrelation, and how they are summarized by two principal components.)

**Figure 3: Correspondence between model and neural representational geometry.

Improvements for AI systems

Based on the findings of this scientific paper, here are specific improvements that could be made to AI systems, along with what those improved systems could achieve:


) Improvement 1: Integrate Neural Representational Geometry into Model Design Constraints (Instead of relying solely on predictive accuracy.)

The paper demonstrates that high predictive accuracy for urban scene appraisal does not guarantee that a model organizes scenes in the same way the human brain does. Furthermore, steering features towards neural geometry to improve prediction fails to raise correspondence with neural structure on unseen data, suggesting this approach overfits rather than recovering a generalizable subspace.

  • The Improved AI System: A multimodal generative architecture where the loss function is augmented not only by standard appraisal prediction error but also by a term penalizing deviations from learned representational geometry (e.g., minimizing the distance in the neural RDM space).

  • What it can do: Create vision models that are not just accurate predictors of human ratings, but are structurally aligned with how humans organize and perceive scenes in a low-level, brain-like manner. This would allow for perceptually consistent representations where internal distances correspond to human perceptual distances.

) Improvement 2: Develop Robust, Noise-Invariant Representation Calibration (Using the Neural Noise Ceiling.)

The paper establishes a clear noise ceiling derived from the agreement between participants' own neural geometries, providing a reproducible benchmark for what structure is consistent across people. Current models are often evaluated against this ceiling without accounting for it properly.

  • The Improved AI System: A representation evaluation framework that explicitly uses the noise ceiling (the lower bound of participant variability) as a calibrated denominator for measuring model correspondence, rather than just reporting raw correlation percentages.

  • What it can do: Provide an objective, truth-grounded metric for judging the quality of learned embeddings. It would allow researchers to definitively state whether a new architecture has captured meaningful neural structure or merely high-frequency noise that happens to correlate with appraisal scores.

) Improvement 3: Implement Task-Set Geometry Invariance (Decoupling Appraisal Goal from Scene Representation.)

The research shows that the underlying neural geometry of a scene does not shift based on the viewer's evaluative goal (e.g., safety vs. beauty). This suggests that models should learn representations invariant to these high-level appraisal prompts.

  • The Improved AI System: A modular model design where the core feature extractor is trained using a cross-task objective that explicitly enforces invariance to different evaluation scales, ensuring the learned embedding captures scene structure rather than goal-specific features.

  • What it can do: Build AI systems for urban analytics (like safety assessment) that are inherently robust to the specific appraisal metric chosen. The system would learn what a street is, independent of whether it's being judged for beauty or safety.

) Improvement 4: Prioritize Low-Level Structural Features over High-Level Semantic/Category Features in Early Layers (Leveraging Hierarchical Correspondence.)

The study found that deeper layers generally matched later neural responses, preserving hierarchical correspondence related to object recognition, even when overall correspondence was low. Furthermore, simple descriptors like Gabor energy matched the best models.

  • The Improved AI System: Architectures that explicitly enforce or prioritize the learning of low-level spatial and structural features (like edges, textures, and local layout) in early transformer blocks before higher semantic abstractions are learned in deeper layers.

  • What it can do: Create faster, more interpretable models that capture the how of scene composition (e.g., density of built structures vs. vegetation) before capturing the what (e.g., specific objects). This would lead to more robust feature extraction for tasks like automated scene parsing or structural analysis in city planning.

Related papers