Evaluation Resolution Confounds Learning-Rule Comparisons in Model-Brain RSA of Early Visual Cortex
q-bio.NC, cs.LG
Submitted: 2026-08-11
Updated: 2026-10-03
Comments: 13 pages, 7 figures
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper investigates a methodological confound in representational similarity analysis (RSA) comparisons of learning rules against brain data at early visual cortex.
Terminology
Summary
The paper investigates a methodological confound in representational similarity analysis (RSA) comparisons of learning rules against brain data at early visual cortex. The central finding is that the common qualitative result—that untrained or locally trained networks rival or beat backpropagation at V1—depends strongly on the resolution at which the network is evaluated.
Specifically, the V1 gap between an untrained network and a backpropagation-trained one widens with evaluation resolution, from −0.001 ± 0.007 at the 32 px training resolution to +0.044 ± 0.006 at 224 px, growing monotonically across six resolutions (n = 5 seeds). The gap at the training resolution is ≈ 0 under the fixed Conv1→V1 mapping and +0.014 under best-layer selection, and the growth with resolution holds under both. Every trained condition (backprop, feedback alignment, predictive coding, STDP) aligns best at or near the 32 px training resolution and falls off as resolution rises: backprop from ρ = 0.065 at 32 px to ρ = 0.031 at 224 px, feedback alignment from 0.020 to 0.012, predictive coding from 0.026 to 0.016, STDP from 0.059 to 0.037. Only the untrained random network rises, from ρ = 0.065 to ρ = 0.076.
The effect holds across all five conditions, in both human fMRI and (directionally, single-seed) macaque electrophysiology, along the full training trajectory, and for two architectures trained at 224 px on other data (an ImageNet ResNet-50 and a Swin-Tiny transformer). The ResNet-50 and the transformer also align best at low resolution despite being trained at 224 px, ruling out train/eval resolution matching as an explanation.
Four candidate mechanisms are tested and none accounts for the effect: (1) train/eval resolution matching, (2) low-level Gabor and pixel structure, (3) the normalization state of the untrained baseline (tested with a 2×2 calibration design holding convolutional weights bit-identical), and (4) convergence of the pooled descriptor toward a global brightness statistic. A fifth experiment locates the effect: repeating the sweep on stimuli first reduced to 32 px and then upsampled caps the image detail at the training resolution while the network still pools over the full number of positions. Across the 64 → 224 px span, the pooled positions increase twelvefold while content stays fixed, and the Random−Backprop gap opens by +0.003 ± 0.001 with content fixed versus +0.030 ± 0.002 with content free to vary. Backprop’s decline is abolished (−0.023 → −0.000, 0/5 and 2/5 seeds). The dependence is carried by image detail above the training resolution rather than by the number of pooled positions.
One control result is stated separately: a single scalar luminance value per image reaches ρ = 0.075 against V1, essentially matching the untrained network’s 0.076. Partialling luminance out halves the untrained network’s alignment (0.076 → 0.038). Luminance similarity orders the conditions exactly as their resolution slopes do, but it is not the carrier of the effect. The paper notes that the first of those numbers concerns the brain data alone and is independent of any model; the second is specific to RSA on globally pooled features, and a fitted readout on the full feature map might place the models higher. What the pair bounds is what this style of comparison can resolve at V1 in this dataset, not model–brain alignment in general.
The one learning effect that holds across resolution is backprop above untrained at a higher area (LOC): backprop−untrained ≈ +0.019, 5/5 seeds at both 32 px and 224 px. IT shows a weaker but consistently positive version of the same thing (+0.015 ± 0.002 at 32 px, +0.005 ± 0.001 at 224 px; 5/5 seeds).
The paper also reports a correction to the author's earlier work: the predictive-coding and STDP conditions had overridden eval with a no-op, leaving their batch-normalization layers in training mode during feature extraction. Repairing this changes those two conditions substantially and leaves the other three unchanged to within 0.0013. The training-dynamics result does not survive repair: with the defect fixed, predictive coding degrades V1 alignment more than backpropagation does rather than less.
The paper's contributions are: (1) identifying and quantifying an evaluation-resolution dependence in model–brain RSA at early visual cortex, general across conditions, two species, the training trajectory, and three architecture families; (2) testing four candidate mechanisms and ruling out all four, three by interventions holding convolutional weights bit-identical; (3) separating the two things evaluation resolution changes at once and locating the dependence on the image-content axis rather than the pooling axis; (4) showing that a single scalar luminance value matches the untrained network's V1 alignment, and that luminance similarity orders conditions exactly as their resolution slopes do without carrying the effect; (5) isolating a learning effect that survives across resolution at LOC, and recommending evaluating alignment at the training resolution and at several others.
The conclusion states: "RSA comparisons of learning rules and architectures at early visual cortex are governed by an evaluation-resolution confound that can manufacture or substantially change the apparent advantage of untrained networks. It holds across conditions, species, training time, and three architecture families. Four candidate mechanisms do not explain it... A fifth experiment locates it: capping image detail at the training resolution while letting the pooled positions grow 12-fold removes about 90% of the effect, so what varies with evaluation resolution is the image detail and not the number of averaged positions. Along the way we find that a single scalar luminance value per image matches the untrained network's V1 alignment here, and that luminance similarity orders the conditions exactly as their resolution slopes do without carrying the effect. That is a caution about what this comparison can resolve at all. In this setting, the learning effect that holds across resolution sits at a higher area (LOC). For this line of work, evaluation resolution is something to control and to report."
Improvements for AI systems
Improvements to AI Systems:
- Resolution-Aware Evaluation Protocol for Model–Brain Alignment Metrics
-
Implement a multi-resolution evaluation harness for any RSA-based comparison of neural network learning rules against brain recordings.
-
The system will automatically compute alignment scores (e.g., Spearman ρ) at the training resolution plus at least three higher resolutions (e.g., 2×, 4×, 8×).
-
It will flag any conclusion that flips or changes sign across resolutions as
resolution-dependent
and suppress claims of superiority for untrained or locally trained models unless the effect persists at all resolutions. -
This prevents false positives in neuroscience-inspired model selection, especially for early visual cortex benchmarks.
- Luminance-Confound Correction for Pooled Feature RSA
-
Add a preprocessing step that partials out a global scalar luminance value (mean pixel intensity per image) from both the model’s pooled feature vectors and the brain’s voxel responses before computing RSA.
-
The system will report both raw and luminance-partialled alignments, and if the raw alignment drops by >50% after partialling, it will label the result as
luminance-dominated
and exclude it from model ranking. -
This makes model–brain comparisons more robust against trivial image statistics that can artificially inflate untrained network performance.
- Batch-Normalization State Audit for Feature Extraction
-
Build an automated checker that verifies all BatchNorm layers are in
evalmode (using running statistics) during any feature extraction for RSA or representational comparisons. -
The system will log a warning and abort the comparison if any BN layer is found in
trainingmode, preventing silent confounds like the one that corrupted predictive-coding and STDP results in the paper. -
This ensures reproducibility and correctness for any learning-rule comparison pipeline.
- Resolution-Content Separator for Image Detail vs. Pooling Effects
-
Develop a diagnostic tool that, given a model and a set of images, decomposes the effect of evaluation resolution into two axes: (a) image detail (high-frequency content) and (b) number of pooled spatial positions.
-
The tool will run two sweeps: one with native resolution increasing, and one with images downsampled to the training resolution then upsampled (capping detail) while pooling positions increase.
-
It will output the contribution of each axis to any alignment change, allowing researchers to identify whether their results are driven by content or by averaging statistics—preventing misinterpretation of resolution effects.
- Cross-Area Robustness Check for Learning Effects
-
Extend any model–brain RSA pipeline to automatically test learning-rule differences (e.g., backprop vs. untrained) not only at early visual cortex (V1) but also at higher areas (e.g., LOC, IT) if such data are available.
-
The system will require that a claimed learning effect replicates across at least two brain areas and across multiple resolutions before being accepted as a genuine training signal.
-
This filters out area-specific artifacts and highlights effects that generalize, as the paper found for backprop at LOC.
- Resolution-Slope Reporting for Model Comparisons
-
Add a standardized output metric: the slope of alignment (ρ) versus log-resolution for each model condition, with confidence intervals across seeds.
-
The system will automatically classify models as
resolution-increasing
(e.g., untrained networks) orresolution-decreasing
(e.g., trained networks) and flag any comparison where the ordering of models changes with resolution. -
This makes the confound visible at a glance and forces researchers to report the full resolution sweep rather than a single point.
What the Improved AI System Can Do:
-
Reliably rank learning rules (backprop, feedback alignment, predictive coding, STDP) against brain data without being misled by evaluation-resolution artifacts.
-
Automatically reject false claims of untrained-network superiority at V1 by requiring multi-resolution consistency and luminance partialling.
-
Detect and fix BN-mode bugs in any feature extraction pipeline, ensuring that reported results are not artifacts of training-mode statistics.
-
Separate the causes of resolution-dependent alignment (image detail vs. pooling count) for any model–brain comparison, enabling targeted interventions.
-
Identify genuine learning effects that survive across brain areas and resolutions, such as backprop’s advantage at LOC, while discarding spurious V1 effects.
-
Produce transparent, reproducible reports that include resolution sweeps, luminance controls, and BN audits, making model–brain alignment studies more rigorous and comparable across labs.
Abstract
Representational similarity analysis (RSA) is increasingly used to ask which learning rules give convolutional networks brain-like representations. Because biologically plausible rules such as feedback alignment, predictive coding and STDP do not scale, studies that include them train small networks on small images (typically 32x32 CIFAR) and then compare them to brain responses modeled at much higher resolution. We find that a common result in this setting, that untrained or locally trained networks rival or beat backpropagation at early visual cortex, depends strongly on the resolution at which the network is evaluated. The V1 gap between an untrained network and a backpropagation-trained one widens from-0.001 +/- 0.007 at the 32px training resolution to +0.044 +/- 0.006 at 224px, growing monotonically across six resolutions (n=5 seeds). It holds in human fMRI and, directionally, in single-seed macaque electrophysiology, along the training trajectory, and for an ImageNet ResNet-50 and a Swin-Tiny transformer trained at 224px. Four candidate mechanisms are tested and none accounts for it: train/eval resolution matching, low-level Gabor and pixel structure, the normalization state of the untrained baseline, and convergence of the pooled descriptor toward a global brightness statistic; three are excluded by interventions holding the convolutional weights bit-identical. A fifth experiment locates the effect: capping image detail at the training resolution while letting the pooled positions grow 12-fold removes about 90% of it, so the dependence is carried by image detail rather than by pooling. Separately, a single scalar luminance value per image reaches rho = 0.075 against the V1 RDM, essentially matching the untrained network's 0.076, which bounds what this style of comparison can resolve. The one learning effect that holds across resolution is backprop above untrained, at LOC.
Sources
- Correspondence of Deep Neural Networks and the Brain for Visual Textures
- Untrained CNNs Exceed Backpropagation in V1 Alignment at High Evaluation Resolution: A Systematic RSA Comparison of Four Learning Rules Against Human fMRI
Related papers
- BrainWave: A Brain Signal Foundation Model for Clinical Applications
- Toward Robust, Reproducible, and Widely Accessible Intracranial Speech Brain-Computer Interfaces: A Comprehensive Narrative Review of Neural Mechanisms, Hardware, Algorithms, Evaluation, Clinical Pathways and Future Directions
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- Emergence of psychopathological computations in large language models
- NeuroAI and Beyond: Bridging Between Advances in Neuroscience and Artificial Intelligence
- Attraction to hierarchical feature memory explains orientation bias