NTH

Evaluation Resolution Confounds Learning-Rule Comparisons in Model-Brain RSA of Early Visual Cortex

AuthorsNils Leutenegger

September 3, 2026 3 min read
Watch on YouTube
The one-line take

The study shows that claims about which vision-learning rules best match the brain can change dramatically depending on the image resolution used for evaluation.

Key results

8,000
CIFAR-10 training subset

Images used to train the small convolutional networks.

−0.001
V1 random-minus-backprop gap at 32 px

Spearman-correlation gap at the training resolution.

+0.044
V1 random-minus-backprop gap at 224 px

Spearman-correlation gap at standard high-resolution evaluation.

+0.003
Content-controlled gap

Gap growth when image detail is capped at 32 pixels while pooled positions increase twelvefold.

0.074
V1 luminance RSA

Spearman correlation between a scalar luminance descriptor and the V1 RDM.

+0.019
LOC backpropagation advantage

Backpropagation-minus-random alignment gap at both 32 and 224 pixels.

What the paper found

This study shows that evaluation resolution can reverse conclusions about which learning rules produce brain-like representations in early visual cortex. A small convolutional network trained for 40 epochs on an 8,000-image CIFAR-10 subset at 32 × 32 pixels was compared using representational similarity analysis against THINGS-fMRI responses to 720 object images, testing random weights, backpropagation, feedback alignment, predictive coding, and STDP. Across six evaluation resolutions from 32 to 224 pixels, the random network’s V1 alignment increased while backpropagation’s decreased: their Spearman-correlation gap shifted from −0.001 ± 0.007 at 32 pixels to +0.044 ± 0.006 at 224 pixels. The pattern persisted after batch-normalization calibration, across human fMRI and directionally in macaque electrophysiology, and in ImageNet-trained ResNet-50 and Swin-Tiny models, ruling out simple train–evaluation resolution matching, Gabor or pixel statistics, and normalization state as sufficient explanations. A content-control experiment downsampled images to 32 pixels before upsampling them, keeping pooled positions increasing twelvefold while image detail stayed fixed; the random–backpropagation gap then grew only +0.003 ± 0.001, versus +0.030 ± 0.002 with unconstrained detail, implicating image content above training resolution rather than pooling alone. A single scalar luminance descriptor reached ρ = 0.074 against V1, essentially matching the untrained network’s 0.075, highlighting how little globally pooled RSA may resolve in this dataset. The robust learning effect appeared at LOC, where backpropagation exceeded random weights by +0.019 at both 32 and 224 pixels. The recommendation is to report RSA at training resolution and across multiple evaluation resolutions.

Original abstract

Representational similarity analysis (RSA) is increasingly used to ask which learning rules give convolutional networks brain-like representations. Because biologically plausible rules such as feedback alignment, predictive coding and STDP do not scale, studies that include them train small networks on small images (typically 32x32 CIFAR) and then compare them to brain responses recorded for naturalistic stimuli modeled at far higher resolution. We find that a common result here -- that untrained or locally trained networks rival or beat backpropagation at early visual cortex -- depends strongly on the resolution at which the network is evaluated. The V1 gap between an untrained and a backprop-trained network widens from -0.001 +/- 0.007 at the 32 px training resolution to +0.044 +/- 0.006 at 224 px, growing monotonically across six resolutions (n = 5 seeds). It holds in human fMRI and, directionally, in single-seed macaque electrophysiology, along the training trajectory, and for an ImageNet ResNet-50 and a Swin-Tiny transformer trained at 224 px. We test four candidate mechanisms and none accounts for it: train/eval resolution matching, low-level Gabor and pixel structure, the normalization state of the untrained baseline, and convergence of the pooled descriptor toward a global brightness statistic. A fifth experiment locates it: capping image detail at the training resolution while letting the pooled positions grow 12-fold removes about 90% of the effect, so the dependence lives on the image-content axis. One control result is worth stating separately: a single scalar luminance value per image reaches rho = 0.074 against V1, matching the untrained network's 0.075, bounding what this comparison can resolve at V1 here. The one learning effect that holds across resolution sits at LOC. Comparisons at early visual cortex must control, and report, the evaluation resolution.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis