Perceptual misalignment of texture representations in convolutional neural networks

arXiv:2604.01341 · cs.CV, q-bio.NC · Submitted 2026-04-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Perceptual misalignment of texture representations in convolutional neural networks".

Tom: The gist The study quantifies perceptual content captured by feature correlations computed for a diverse pool of Convolutional Neural Networks and compares it to their perceptual alignment with the mammalian…

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So we're diving into this paper today, "Perceptual misalignment of texture representations in convolutional neural networks." It sounds super technical, but basically, they're looking at how well these deep learning models actually capture what we perceive as texture.

Jane: Exactly. The title itself points to a mismatch—a perceptual misalignment—between the features the computer sees and what a human actually feels when they look at that surface. It’s about checking if these CNNs are really modeling texture, or just some other kind of pattern recognition that looks good on paper.

Lu: They take this idea from Julesz back to local correlations, which is a big concept in vision science, and then they apply it using Gram matrices from CNN activations. That's the core method they're testing here.

Meng: So, the main question they’re asking is whether these texture representations spontaneously align with the textures’ actual perceptual content. It seems like they are testing if standard object recognition CNNs actually capture texture perception well at all.

Tom: Right, and according to page two of this paper, their surprising finding is that there's no connection between conventional measures of CNN quality as a model for the visual system and how well those textures align with human texture perception when measured by Brain-Score <ref:2604.01341#pg1>.

Jane: That’s a big statement. It suggests that the features CNNs learn for object recognition aren't necessarily what drives how we perceive texture, which is pretty counterintuitive since these networks are so common now.

Lu: They conclude that texture perception involves mechanisms different from those commonly modeled using CNNs trained on object recognition, possibly depending on how contextual information gets integrated into the system.

Meng: From an engineering standpoint, this means if we rely only on standard object recognition CNNs to model texture, we're missing something crucial about how humans process visual information.

Tom: And they do suggest some improvements for the method itself, which is interesting because it’s not just saying "this doesn't work." They propose ways to make the texture synthesis and representation quality better.

Jane: They suggest weighting layer losses based on feature map dimensions, which they argue should be done by dividing each layer-specific Gram loss by four times the number of spatial samples squared. That’s a concrete way to prevent one big layer from dominating the synthesis process.

Lu: And they also point out that for most networks, perceptual quality in these Gram representations tends to grow monotonically with the depth of the layers analyzed, which is consistent with ideas about texture processing being distributed along intermediate areas of the ventral visual stream like V2 and V4.

Meng: That monotonic growth is a good sign because it aligns with where we expect texture processing to happen in the brain, but they also found that this Gram representation isn't perfect at recovering the full class structure determined by human labels.

Title and authors: Tom: That means even when we get the general idea of texture right, we still can't perfectly reconstruct every specific type of texture based on what humans label them as. The clustering MI only reaches about half of the theoretical maximum according to page two <ref:2604.01341#pg1>.

Jane: So, while the process shows a trend moving deeper into the network layers, it’s not achieving perfect classification fidelity for textures. It's an improvement in quality, but it's still fundamentally misaligned with human judgment on texture categories.

Lu: The paper implies that if we want to build better methods for texture analysis and synthesis that truly align with human perception, we need to look beyond standard CNNs trained purely for object recognition tasks.

Meng: That points toward needing more specialized training, maybe unsupervised or self-supervised learning techniques instead of just relying on what’s learned from ImageNet-1K <ref:2604.01341#pg1>.

Tom: Right. And they suggest that future architecture design should focus more on perceptual alignment metrics rather than just standard object recognition performance metrics to see better results.

Jane: They also touched upon how texture bias emerges primarily at the level of the decision head, which is a weird place for this kind of misalignment to show up, but it’s something they flag as an area for future study.

Lu: And because textures are distributed along intermediate areas of the ventral visual stream, synthesizing textures using activations from those specific layers could lead to outputs that are more aligned with human perception.

Meng: So, we're looking at a path forward where we focus on how the features themselves are structured across different network depths, not just how well the final classification head performs.

Tom: We’re wrapping up this discussion on "Perceptual misalignment of texture representations in convolutional neural networks." It’s clear that while Gram matrices are useful tools for modeling correlations, they don't automatically translate into perceptual accuracy when compared to human judgments.

Jane: The main implication is that the functional features learned by object recognition AI aren't the same ones we use to understand how humans perceive texture, which makes our current AI models a bit of a mismatch for this specific task.

Lu: The paper lays out a path: use deeper layers for better structure recovery and perhaps look at different learning paradigms entirely to bridge that gap between computer vision and human visual neuroscience.

Meng: For practical applications, this means any texture synthesis tool we build needs to be carefully tuned not just for accuracy on a benchmark, but for how those features map onto the visual stream areas we discussed.

Tom: So, let’s leave it there for now on this paper. It makes us rethink what "quality" means when we evaluate a CNN as a model of human vision. Next up, we're looking at how digital personas are doing in approximating human survey findings.

Jane: We’ll be right back after this break to talk about that paper and see if those AI substitutes for people can actually get the answers they claim to provide.

The paper's summary: Tom: So, to recap this paper, they’re basically showing that when we use those Gram matrices from CNNs to model textures, they don't actually line up with how humans judge texture quality.

Jane: Right. It’s about checking if the way a computer sees texture matches what we perceive as good or bad texture. The main finding is there's no connection between how good the network’s internal math looks and whether it captures human perception of those textures at all.

Lu: What they found is that even though these Gram representations are used for synthesis, they aren't doing a great job recovering the whole category structure humans assign to those textures. It’s only hitting about half of what the theoretical maximum clustering score should be.

Meng: So, it means these standard object recognition models aren't actually learning the specific visual rules that drive human texture perception in a way we can easily measure. That’s a practical concern for anyone building tools that need to generate realistic surfaces.

Tom: Exactly. And they point out something interesting about the depth of the layers: as you go deeper into the network, the perceptual quality of those Gram representations tends to get better, which makes sense if we think about how texture processing happens in our visual system.

Jane: That trend is consistent with ideas that suggest texture information is spread out across different areas of the brain, like V2 and V4. It’s not just a single spot in the network that matters for texture representation.

Lu: They suggest this deeper layer growth might be related to how the features are distributed along those visual streams. It hints at where we need to look if we want better models for texture synthesis, maybe targeting those intermediate areas of processing instead of just whatever standard layers they pick.

Meng: From an engineering standpoint, if we’re going to use these CNNs for anything more than simple object recognition, this tells us that simply training a bigger or deeper network won't automatically solve the texture modeling problem.

Tom: It means we might need to look at different types of learning altogether, like unsupervised methods, instead of just relying on what those standard object recognition models were trained on.

Jane: That’s the big implication for anyone working in computer vision right now—we need to move beyond just optimizing classification accuracy and start thinking about how features map onto human sensory experience.

Lu: It opens up a lot of creative possibilities, suggesting that maybe we should design architectures specifically to capture those kinds of local correlations in a way that’s perceptually meaningful, not just statistically optimized for recognition.

Meng: So the practical step is shifting focus from general object performance to alignment with specific perceptual targets when you’re designing your models.

Tom: Right. This whole discussion on texture representation quality leads us into thinking about how digital personas are handling human survey findings in our next segment.

The paper's improvements: Tom: So, after showing that there's this disconnect between CNN features and human texture perception, the authors suggest some ways to actually improve those representations when we synthesize textures.

Jane: They’re proposing weighting the loss from different layers differently so that one big layer doesn't completely dominate what the AI learns during synthesis. It sounds like a way to keep all parts of the feature map in balance.

Lu: That makes sense because they noticed that for most networks, the quality of those Gram representations just keeps growing as you go deeper into the layers analyzed. So they think focusing on those deeper layers might give you a clearer structure for clustering.

Meng: So, if we’re building a texture generator, this means we shouldn't just pick one layer to pull features from; we need to consider the whole depth of the network for better results.

Tom: And they also link that deeper growth back to visual neuroscience, suggesting it lines up with how texture is processed in areas like V2 and V4 in our brains. That’s a big connection there.

Jane: It sounds like they are trying to make the computer's internal representation more similar to the actual way our eyes and brains work when we look at surfaces.

Lu: They’re suggesting that instead of just using standard layers, we should try to target those specific visual stream areas for our synthesis models. That could actually lead to outputs that feel more perceptually correct.

Meng: So the practical implication is that if we want better texture synthesis, we need to stop looking at the network as a black box and start thinking about which parts of the network correspond to which parts of human vision.

Tom: Exactly. And they also point out another thing: they found that this Gram method isn't perfect at capturing all the different types of textures humans label. It’s only getting halfway to that full class structure fidelity.

Jane: That means even if the process is getting deeper and better, it still doesn't perfectly match the entire spectrum of human texture categories we use for labeling.

Lu: It shows that this method is a step in the right direction, but it isn't the final answer for modeling perceptual reality. It’s a tool that shows where we need to go next to get closer to what humans experience.

Meng: So, the takeaway here is that designing texture tools needs to move beyond just standard CNN metrics and start incorporating these kinds of perceptual alignment ideas into the training process itself.

Tom: Right. This whole discussion on improving texture representation quality points us toward rethinking how we train AI for visual tasks altogether. That brings us right back to those digital personas we were talking about earlier, and whether they can actually get the answers they claim to provide.

Conclusion: Tom: So to wrap up this whole thing, we’re looking at the paper "Perceptual misalignment of texture representations in convolutional neural networks." Basically, they proved that standard CNN features don't perfectly match how we perceive textures.

Jane: That’s right. The main message is that conventional measures of a network's quality don't tell us anything about its alignment with human texture perception at all. It’s a bit surprising since these networks are everywhere now.

Lu: What they show is that while deeper layers in the network might have better internal structure, the resulting Gram representations still aren't perfectly recovering the full set of textures humans assign to those images. They only hit about half of what the theoretical maximum clustering score should be.

Meng: So, for anyone building a texture synthesis tool, this means you can't just optimize for general CNN quality metrics and expect it to produce perceptually accurate results on its own.

Tom: Exactly. And they point out that because texture processing seems distributed across different areas of the visual stream, like V2 and V4, we should be looking at how we use those deeper layers instead of just stopping at a certain point in the network.

Jane: It’s about aligning the computer's internal math with how our brains actually work when we look at a surface. That’s a pretty deep idea to take away from this research.

Lu: It really opens up avenues for creative design, suggesting that maybe we need to look at different learning methods entirely, like unsupervised learning, to get those better perceptual results.

Meng: From an engineering side, it means future architectures might need to incorporate self-supervised learning techniques instead of just relying on what's learned from massive object recognition datasets.

Tom: It shifts the focus from just making a model that classifies objects well to making a model that understands visual perception better overall.

Jane: It really makes you wonder what other aspects of AI we need to look at when we try to build systems that interact with the real world in ways they look more human.

Lu: It points toward building models that are specifically tuned for those intermediate processing areas, trying to capture texture information where it actually gets processed in the visual system.

Meng: So, the implication is that we need a new set of metrics for evaluating these AI systems—metrics that focus on perceptual alignment rather than just standard recognition scores.

Tom: Exactly. We’re leaving this paper with the idea that texture modeling is more complex than just correlating features; it requires understanding how those features are distributed across the visual system.

Jane: It’s a reminder that even powerful tools like CNNs have limitations when it comes to capturing subtle, human-like sensory details like texture.

Lu: We still have a lot of room to explore how we can bridge that gap between current computer vision methods and true perceptual modeling.

Meng: And once we figure out better ways to model texture, it could influence how we design interfaces or even how robots interact with surfaces in the real world.

Tom: Right. Speaking of things interacting with the real world, let’s shift gears completely to talk about those digital personas that are trying to mimic human survey results.

Ludovica de Paolis, Fabio Anselmi, Alessio Ansuini, Eugenio Piasini

International School for Advanced Studies (SISSA) · Department of Mathematics, Informatics and Geosciences, Università degli Studi di Trieste · Department of Data Engineering, Area Science Park

cs.CV, q-bio.NC

Submitted: 2026-04-01

Updated: 2026-05-18

Code: https://github.com/ludovicadepaolis01/perceptual_misalignment

Importance score: 68/100

The gist: The gist The study quantifies perceptual content captured by feature correlations computed for a diverse pool of Convolutional Neural Networks and compares it to their perceptual alignment with the

Key concepts

Gram Matrix
A mathematical tool used to model textures by calculating the inner product of feature maps from different layers of a CNN. It captures the spatial correlations between features, which are believed to encode the essential elements needed for texture perception.
Texture Synthesis
The process where an image is iteratively modified using gradient descent to match the statistical structure (specifically, the Gram matrices) of a target texture image. This demonstrates whether a CNN's learned features can actually generate perceptually convincing textures.
Perceptual Alignment
The degree to which the statistical representations learned by a CNN (like its Gram matrices) correspond to what humans actually perceive as texture quality. The study found that conventional measures of CNN performance do not correlate with this human perceptual alignment.
Ventral Visual Stream
A region in the visual system, including areas like V2 and V4, believed to be involved in processing complex visual information such as textures. The paper suggests that texture processing might be distributed across these intermediate areas of the brain.

Terminology

Summary

The gist The study quantifies perceptual content captured by feature correlations computed for a diverse pool of Convolutional Neural Networks and compares it to their perceptual alignment with the mammalian visual system, revealing no connection between conventional measures of CNN quality and human texture perception

Introduction

Visual textures hold a special status in vision neuroscience as stimuli that afford a nontrivial degree of controllability and complexity, unlike natural images or more traditional parametric stimuli Because of this, multiple mathematical frameworks for modeling and generating textures have been established, with considerable success One particularly influential strand of theoretical, computational and empirical work stems from the seminal view of Julesz (Julesz, 1981; Julesz, 1962; Caelli & Julesz, 1978), who proposed that texture perception is controlled by local correlations of visual features within an image This view was extended by later work that better formalized the notion of textures as statistical objects (Zhu et al., 2000) In recent years, following the rise in popularity of deep convolutional neural networks (CNNs) as tools in computer vision, Gatys et al. proposed a machine learning-based operationalization of Julesz’s idea (Gatys et al., 2015) The Gatys model effectively defined a texture ensemble based on the Gram matrices formed by spatial correlations between intermediate-layer activations of a CNN trained for object recognition Texture synthesis is performed by optimizing the image pixels via gradient descent to match the statistical structure of an exemplar texture

Methods

The study employed a method where textures are modeled using feature representations extracted from a CNN, which are shown to support the generation of perceptually compelling texture samples

This Gram representation are thought to capture the elements of the input image that are important for texture perception

The process involves several key steps:

  1. Feature Extraction and Gram Matrix Computation: Given a natural texture image x, the algorithm starts by feeding x as input to the network and extracting the corresponding activation maps from a predetermined set of layers For each selected layer l, the feature activations F l are used to compute a Gram matrix Gl, defined as the inner product between vector representations of the feature maps: G l ij = X m F l imF l jm, where the index m runs over spatial samples of the feature maps, and i and j label feature channels at layer l

  2. Texture Synthesis: Texture synthesis is performed by iteratively updating the pixels of an initially random image xˆ, where pixel values are drawn from a Gaussian distribution xˆ ∼ N (0, 1) This process leads to match the Gram matrices of xˆ to those of the target image x by minimizing the Gram loss, defined as a weighted sum of squared L2 distances between corresponding Gram matrices across layers L(ˆx, x) = X l 1/4MlN2 l Gˆl − G l 2 (2)

  3. CNN Layer Selection: The analysis used a set of popular CNNs trained for object classification with ImageNet-1K, and feature maps were extracted from five layers from each architecture, selected in such a way that they are located at comparable relative depth across architectures

  4. Stimulus Dataset: The Describable Textures Dataset (DTD) is the input dataset of choice, consisting of 5640 images of naturalistic visual textures collected from the Internet, spanning a wide range of materials and surface appearances

Results

The analysis focused on two main comparisons: the quality of texture representations against human perception and the alignment with models for object recognition

Our results show that there is no correlation between the goodness of the Gram texture representation – with respect to the clusters formed by human-annotated labels – and Brain-Score

The study found that there is no correlation between conventional measures of CNN quality as a model of the visual system and its alignment with human texture perception Specifically, they found no correlation between the goodness of the Gram texture representation – with respect to the clusters formed by human-annotated labels – and Brain-Score

The perceptual quality of the Gram representations tends, for most networks, to grow monotonically with the depth of the layers analyzed

The perceptual quality of the Gram representations tends, for most networks, to grow monotonically with the depth of the layers analyzed This is generally compatible with recent results suggesting that texture processing is distributed along intermediate areas of the ventral visual stream including V2 and V4 (Okazawa et al., 2017; Ziemba et al., 2024) and possibly IT (Zhivago & Arun, 2014)

Furthermore, the Gram representation is unable to recover the full class structure of the data, as determined by the human-assigned labels Quantitatively, this is seen by the value of the clustering MI reaching only about half of the theoretical maximum (Figure 2)

Discussion

The analysis leads to three main conclusions regarding the relationship between texture representation quality and model quality

First, the perceptual quality of the Gram representations tends, for most networks, to grow monotonically with the depth of the layers analyzed

This trend is generally compatible with recent results suggesting that texture processing is distributed along intermediate areas of the ventral visual stream including V2 and V4 (Okazawa et al., 2017; Ziemba et al., 2024) and possibly IT (Zhivago & Arun, 2014)

Second, the Gram representation is unable to recover the full class structure of the data, as determined by the human-assigned labels

This highlights that, despite its popularity, the Gram representation is far from perfect

Third, the quality of the CNNs as implicit models for texture perception is not correlated with their quality as models of the ventral stream in an object recognition setting

This result is surprising, because one could have expected that the functional features captured by Brain-Score are general enough to at least partly include those that are necessary to support human-like texture perception The conclusion is that in order to design methods of texture analysis and synthesis that are better aligned with human perception we need to look beyond standard CNNs trained for object recognition, for instance by incorporating unsupervised or self-supervised learning (Matteucci et al.

Improvements for AI systems

  1. Improve texture representation quality by weighting layer losses based on feature map dimensions: Therefore, to make sure that Gram losses coming from different layers have the same weight in our total loss we should divide each layer-specific Gram loss by 4MN squared. This ensures that the model is not overly dominated by a single, high-dimensional layer during synthesis.

  2. Enhance texture class recovery fidelity using deeper feature maps: The analysis shows that the clearest structure emerges in the last three layers, suggesting that using features from deeper layers will yield better results for clustering, as MI grows along with the depth of the extracted features.

  3. Develop more perceptually accurate texture synthesis by targeting specific visual stream areas: Since texture processing is distributed along intermediate areas of the ventral visual stream including V2 and V4, synthesizing textures using activations from these layers could produce outputs more aligned with human perception.

  4. Guide architecture design based on perceptual alignment rather than just object recognition metrics: The finding that the quality of the CNNs as implicit models for texture perception is not correlated with their quality as models of the ventral stream in an object recognition setting suggests that future architectures should incorporate unsupervised or self-supervised learning (Matteucci et al., 2024), multimodality (Radford et al., 2021) or different architectures, such as those with attention mechanisms.

  5. Incorporate texture bias mitigation during training: Since texture bias emerges primarily at the level of the decision head, AI systems should be trained to decouple shape and texture information, potentially by using methods that address this texture bias to ensure representations are more generalizable for texture processing.

Sources

Related papers