NTH

Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws

AuthorsLiu Ziyin, Yizhou Xu, Tomaso Poggio, Isaac Chuang

August 23, 2026 3 min read
Watch on YouTube
The one-line take

This work argues that plateaus, sudden learning, and power-law scaling can all emerge from one simple mathematical model of how neural-network modes turn on.

Key results

17
Unified model instances

Previously separate solvable proxy models represented as neural quadratic forms

0.01
Small initialization scale

Scale used in close NQF approximation experiments

0.2
Large initialization scale

Scale used to test the breakdown of the local approximation

1000
Feature-wise synthetic setup

Number of synthetic samples used to verify the predicted power-law slope

4000
MNIST odd-versus-even task

Number of MNIST samples used for architecture-matching experiments

What the paper found

“Neural Quadratic Forms” proposes a single local model for sudden learning and smooth training-time scaling laws. The key result is that permutation symmetry among interchangeable units, smoothness, and zero gradient at zero force a near-initialization expansion of the form f_x(W)−f_x(0)=μᵀg(x)+Tr[WWᵀA(x)]+O(||W||³), where architecture-specific details are compressed into the structure matrix A(x). Two-layer MLPs, CNNs, mixture-of-experts layers, attention heads, matrix-sensing models, and 17 previously separate solvable proxies become instances of this neural quadratic form. For multi-head attention, the leading term depends on value and readout weights, while query and key weights first appear at quartic order, a result relevant to transformer systems such as ChatGPT but not presented as a ChatGPT benchmark. Training closes on the low-dimensional order parameters M=WWᵀ and μ; under a shared eigenbasis for the data matrices, the feature eigenvalues follow a generalized Lotka–Volterra equation. Each mode activates at a characteristic time proportional to (1/ζ)ln(1/ε), so smaller initialization separates learning events into plateaus, while power-law spectra aggregate them into predictable power-law loss curves. Experiments show close agreement across gradient descent, Polyak momentum, and Adam at initialization scales σ=0.01 and σ=0.2, use a synthetic feature-wise setup with m=1000, and reproduce matching dynamics on a 4000-sample MNIST odd-versus-even task. The theory is limited by small initialization, smooth activations, and the assumption of power-law spectra.

Original abstract

Neural networks trained by gradient descent on a smooth cost function can nevertheless learn in steps: the cost holds on long plateaus and then drops abruptly. Meanwhile, training losses instead follow smooth power laws. Variants of both behaviors occur in architectures with very different microscopic structures, which is the signature of a few relevant collective variables. We show that a symmetry fixes what those variables are: a network layer is a sum over interchangeable units, so relabeling the units leaves it unchanged; given smoothness and the condition that a unit's gradient vanish at the origin, symmetry then enforces a universal leading form for the expansion about the near-zero weights present at the start of training, the quadratic $\Tr[WW^{\top}A(x)]$, in which every architectural detail is confined to a single ``structure matrix" $A(x)$ that we compute for each architecture. Perceptrons, attention layers, mixtures of experts, and convolutions become one model at different $A$. Its training dynamics then close on the ``order parameter" $M=WW^{\top}$ and, whenever the data matrices share an eigenbasis, reduce to a Lotka--Volterra equation whose modes switch on one after another. The smaller the initial weights, the further apart the switch-on times, and the plateaus appear as a singular limit of a smooth flow; when many modes are unresolved the same events merge into a power law in training time whose exponent the theory predicts. We confirm both numerically across training methods and architectures.

Read the original paper

More in Neural Networks

Browse all 22 papers →
02Neural Network

Retrieving Individual Stems from Music Mixtures with Slot Embeddings

David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein

Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.

Read analysis
03Neural Network

The Linear Representation Hypothesis Needs a Group Action

Louie Hong Yao, Yuhao Li, Shengchao Liu

This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.

Read analysis