Untrained CNNs Exceed Backpropagation in V1 Alignment at High Evaluation Resolution: A Systematic RSA Comparison of Four Learning Rules Against Human fMRI

summary

Video file (mp4)

The gist

The gist: Early visual alignment is architecturedriven, learning rules differentiate only at intermediate areas, and all rules converge at the highest levels of the hierarchy<ref:2604.16875#pg2>

In short

The study compared four learning rules—backpropagation, feedback alignment, predictive coding, and STDP—on identical convolutional neural networks against human visual cortex data using RSA. Results show early visual representation alignment is dominated by the network's architecture in areas like V1/V2. Learning rules only show differences at intermediate layers; all methods converge at higher levels (IT), meaning architecture matters more than the learning rule for early feature extraction.

Key concepts

Representational Similarity Analysis (RSA)
RSA is a statistical method used to compare the similarity between two different brain data sets. In this study, it was used to measure how well the internal representations learned by different AI models matched the patterns found in human fMRI scans of the visual cortex.
Architecturedriven Alignment
This concept means that the structure and design of a neural network (the architecture) are more important than the specific way it learns (the learning rule) when determining how its internal features align with human brain structures. The study found that early layers are strongly dictated by the architecture.
V1/V2 vs. IT Hierarchy
The visual cortex is organized hierarchically, moving from simple feature extraction areas like V1 and V2 to higher-level association areas like IT. The research demonstrated that different learning rules affect alignment differently: some rules perform better at early layers (V1/V2), while others converge at the highest level (IT).
Untrained Random-Weights Baseline
This control group consists of a network with random initial weights, meaning it has not been trained. Comparing trained models to this baseline reveals the influence of the architecture itself. The baseline often shows higher alignment in early visual areas than even the best-trained learning rules.

Terminology used across episodes

This episode discusses

The paper

Untrained CNNs Exceed Backpropagation in V1 Alignment at High Evaluation Resolution: A Systematic RSA Comparison of Four Learning Rules Against Human fMRI · Read on arXiv

Nils Leutenegger

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Untrained CNNs Exceed Backpropagation in V1 Alignment at High Evaluation Resolution".

Jane: The gist: Early visual alignment is architecturedriven, learning rules differentiate only at intermediate areas,

Tom: First, who's behind it and why it matters.

Title and authors: Jane: So, we're talking about this comparison between backpropagation, feedback alignment, predictive coding, and spike-timing-dependent plasticity in the paper "Untrained CNNs Exceed Backpropagation in V1 Alignment at High Evaluation Resolution: A Systematic RSA Comparison of Four Learning Rules Against Human fMRI." We've established that the architecture matters early on. Now we look at what the full study actually found.

Tom: The summary of this paper shows a sharp pattern: all five learning rules end up converging at the highest levels of the network hierarchy, like in IT. But they show clear differences happening in intermediate areas.

Meng: So, if we think about practical application for building these kinds of AI systems, what does that convergence actually mean for us when we design models?

Lu: It suggests that while the early stages are dictated by structure, the higher-level abstract representations seem to be robust enough to be shaped regardless of whether you use BP or STDP at those deeper points.

Jane: But they also pointed out that feedback alignment consistently gave the lowest alignment at V1 and V2, which is a bit surprising when you’re looking for biologically plausible models.

Tom: It really highlights that different learning rules have distinct effects depending on where you are in the visual processing pipeline, so we can't just treat all learning algorithms as equally good everywhere.

The paper's summary: Jane: To summarize the core of this paper, they systematically applied backpropagation, feedback alignment, predictive coding, and STDP to identical convolutional architectures and compared their internal representations against human brain data from the THINGS-fMRI dataset using Representational Similarity Analysis.

Tom: They used two hundred twenty-four by two hundred twenty-four resolution for all stimuli and averaged the results across five random seeds to keep things consistent <ref:2604.16875#pg1>. The central question they tackle is whether the learning rule dictates how well the network's internal features align with those of the human visual cortex.

Meng: When we look at their numbers, they found that at V1/V2, predictive coding and STDP outperformed backpropagation in terms of alignment, with STDP hitting a correlation score of zero point zero six four compared to backpropagation's zero point zero three four at V1 under specific conditions.

Lu: That’s a concrete comparison showing that for low-level visual features, these local updates based on prediction error or timing are more effective than the standard backpropagation approach in mimicking brain structure.

Tom: And they also found that feedback alignment consistently produced the lowest brain alignment scores at V1 and V2, which suggests that random feedback filters can actually interfere with what the convolutional architecture is trying to build early on.

Jane: So, it’s a mixed bag: some rules are better at low levels, others hurt them right from the start in those initial visual areas.

The paper's improvements: Tom: The authors point out a few things they think we should consider for future work or for understanding these results better. They mention that their setup uses a small CNN architecture, which limits its overall representational capacity.

Jane: That limitation is important because it might explain why all five learning rules eventually converge at the highest level, suggesting that this convergence could just be a sign of limited model size rather than a universal property of deep learning.

Meng: From an engineering standpoint, they also noted that applying STDP to static images via Poisson spike trains discards some of the temporal dynamics that STDP is usually known for exploiting in actual biological systems. That’s a limitation in their specific experiment setup.

Lu: And they mention a different observation: the best layer-per-ROI analysis showed that the fully connected layer one sometimes produces higher V1 alignment than the first convolutional layer, which complicates how we map these layers onto specific brain regions anatomically.

Tom: So, to summarize their suggested improvements, they are really prompting us to think about using larger architectures or perhaps looking at more complex temporal dynamics if we want to model these systems more accurately.

Conclusion: Jane: Wrapping up this paper on "Untrained CNNs Exceed Backpropagation in V1 Alignment at High Evaluation Resolution: A Systematic RSA Comparison of Four Learning Rules Against Human fMRI," the authors conclude that architecture is dominant early on, and convergence happens later.

Tom: They emphasize that for small-scale setups like this, the search for biologically plausible learning rules should focus more on preserving those architectural inductive biases in the early areas, like V1 and V2.

Meng: So practically speaking, if we’re designing a system inspired by brain function, we should prioritize keeping those structural priors—the local connectivity and weight sharing—over just tweaking the weight update rule for early features.

Lu: And they point to the convergence at IT as a possible indicator of model capacity limits; it suggests that once you get to those high-level abstract representations, the specific learning rule might not matter much anymore.

Jane: So, we see that architecture dictates early visual representations, while higher areas converge regardless of the specific learning algorithm used. That’s what this study shows us about how different rules behave in a CNN context.

More episodes

← Home