Kernel Renormalization in Bayesian Deep Neural Networks: the Equivalent Wishart Ansatz in the Proportional Regime
AuthorsPaolo Baglioni, Christian Keup, Vincenzo Zimbardo, Rosalba Pacelli, Alessandro Vezzani, Raffaella Burioni, Pietro Rotondo
Resources
This paper develops a new theory for how wide Bayesian neural networks generalize, showing that finite-width effects can be captured with a renormalized kernel model that matches experiments reasonably well.
Key results
The theory was validated on deep neural networks with depths up to L ≈ 10.
What the paper found
This paper, from INFN and the University of Parma, develops an approximate statistical-mechanics theory for Bayesian deep neural networks in the proportional regime where sample size P and width N grow together. The key novelty is the Equivalent Wishart Ansatz, which treats the hierarchical empirical kernels of multilayer perceptrons as if their dominant finite-width fluctuations were Wishart-distributed, collapsing an intractable matrix-valued problem into at most L scalar order parameters. That reduction yields a large-deviation effective action for the posterior and a renormalized neural-network Gaussian process kernel, so learning curves can be predicted from a low-dimensional theory rather than P×P kernel self-consistency equations. The authors extend the construction to nonzero-mean activations such as ReLU via a noncentral Wishart analysis, and to convolutional networks through a stacked kernel renormalization scheme that captures patch–patch correlations. They validate the theory against Bayesian posterior sampling with Langevin Monte Carlo, MALA, pCN, and NUTS HMC on MNIST, CIFAR-10, Gaussian teacher-student data, and networks with depths up to L≈10 and P≈10^3, finding excellent agreement except for two systematic breakdowns at large depth and load α=P/N, including a newly reported metastable posterior transition around L>5. In the mean-field/μP parametrization, the theory shows that improved generalization is driven mainly by variance suppression rather than bias change, while the effective kernel remains only globally rescaled. Overall, the result reframes finite-width deep learning as a renormalized-kernel problem governed by sample-to-width scaling, not just lazy versus rich training.
Original abstract
The scaling limit where both the size of the training set $P$ and the width $N$ of a deep neural network grow at the same rate, the so-called proportional-width regime, has been intensely studied for shallow, single-hidden-layer networks. However, extending these non-perturbative results from shallow architectures to deep non-linear networks has proven very challenging. Here we present an effective approximate approach to predict the generalization performance of Bayesian multi-layer perceptrons (MLPs) of fixed depth $L$ on arbitrary high-dimensional data. We propose an equivalent Wishart Ansatz to capture the dominant stochastic fluctuations of the hierarchical empirical kernels of MLPs. This allows us to perform a large deviation analysis for the partition function of MLPs in the proportional limit, expressed in terms of a renormalized NNGP kernel. In this description, even strong representation learning in the proportional limit is encoded in at most $L$ scalar order parameters, determined self-consistently. Extending the approach to convolutional architectures (CNNs), we identify a hierarchical local kernel renormalization mechanism, which allows to quantify more complex data-dependent transformations of the large-width kernel in CNNs due to finite-width effects. We test our effective theory against sampling experiments from the Bayesian posterior of finite deep neural networks with depths $L \sim O(10)$ and $P\sim O(10^3)$ on classic benchmark datasets, finding overall very good agreement together with two distinct types of systematic deviations.
Read the original paperMore in Neural Networks
Browse all 22 papers →End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery
Akira Ito, Takayuki Miura, Yosuke Todo
A new query-efficient technique makes it possible to steal the parameters of small black-box neural networks using only their predicted labels.
Retrieving Individual Stems from Music Mixtures with Slot Embeddings
David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein
Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.
The Linear Representation Hypothesis Needs a Group Action
Louie Hong Yao, Yuhao Li, Shengchao Liu
This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.