Evaluation Resolution Confounds Learning-Rule Comparisons in Model-Brain RSA of Early Visual Cortex
AuthorsNils Leutenegger
Resources
The study shows that claims about which vision-learning rules best match the brain can change dramatically depending on the image resolution used for evaluation.
Key results
Images used to train the small convolutional networks.
Spearman-correlation gap at the training resolution.
Spearman-correlation gap at standard high-resolution evaluation.
Gap growth when image detail is capped at 32 pixels while pooled positions increase twelvefold.
Spearman correlation between a scalar luminance descriptor and the V1 RDM.
Backpropagation-minus-random alignment gap at both 32 and 224 pixels.
What the paper found
This study shows that evaluation resolution can reverse conclusions about which learning rules produce brain-like representations in early visual cortex. A small convolutional network trained for 40 epochs on an 8,000-image CIFAR-10 subset at 32 × 32 pixels was compared using representational similarity analysis against THINGS-fMRI responses to 720 object images, testing random weights, backpropagation, feedback alignment, predictive coding, and STDP. Across six evaluation resolutions from 32 to 224 pixels, the random network’s V1 alignment increased while backpropagation’s decreased: their Spearman-correlation gap shifted from −0.001 ± 0.007 at 32 pixels to +0.044 ± 0.006 at 224 pixels. The pattern persisted after batch-normalization calibration, across human fMRI and directionally in macaque electrophysiology, and in ImageNet-trained ResNet-50 and Swin-Tiny models, ruling out simple train–evaluation resolution matching, Gabor or pixel statistics, and normalization state as sufficient explanations. A content-control experiment downsampled images to 32 pixels before upsampling them, keeping pooled positions increasing twelvefold while image detail stayed fixed; the random–backpropagation gap then grew only +0.003 ± 0.001, versus +0.030 ± 0.002 with unconstrained detail, implicating image content above training resolution rather than pooling alone. A single scalar luminance descriptor reached ρ = 0.074 against V1, essentially matching the untrained network’s 0.075, highlighting how little globally pooled RSA may resolve in this dataset. The robust learning effect appeared at LOC, where backpropagation exceeded random weights by +0.019 at both 32 and 224 pixels. The recommendation is to report RSA at training resolution and across multiple evaluation resolutions.
Original abstract
Representational similarity analysis (RSA) is increasingly used to ask which learning rules give convolutional networks brain-like representations. Because biologically plausible rules such as feedback alignment, predictive coding and STDP do not scale, studies that include them train small networks on small images (typically 32x32 CIFAR) and then compare them to brain responses recorded for naturalistic stimuli modeled at far higher resolution. We find that a common result here -- that untrained or locally trained networks rival or beat backpropagation at early visual cortex -- depends strongly on the resolution at which the network is evaluated. The V1 gap between an untrained and a backprop-trained network widens from -0.001 +/- 0.007 at the 32 px training resolution to +0.044 +/- 0.006 at 224 px, growing monotonically across six resolutions (n = 5 seeds). It holds in human fMRI and, directionally, in single-seed macaque electrophysiology, along the training trajectory, and for an ImageNet ResNet-50 and a Swin-Tiny transformer trained at 224 px. We test four candidate mechanisms and none accounts for it: train/eval resolution matching, low-level Gabor and pixel structure, the normalization state of the untrained baseline, and convergence of the pooled descriptor toward a global brightness statistic. A fifth experiment locates it: capping image detail at the training resolution while letting the pooled positions grow 12-fold removes about 90% of the effect, so the dependence lives on the image-content axis. One control result is worth stating separately: a single scalar luminance value per image reaches rho = 0.074 against V1, matching the untrained network's 0.075, bounding what this comparison can resolve at V1 here. The one learning effect that holds across resolution sits at LOC. Comparisons at early visual cortex must control, and report, the evaluation resolution.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.