NTH

The Implicit Bias of Depth: From Neural Collapse to Softmax Codes

AuthorsConnall Garrod, Jonathan P. Keating, Christos Thrampoulidis

May 29, 2026 2 min read
Watch on YouTube
The one-line take

This paper explains how depth alone can steer neural networks away from neural collapse and toward low-rank softmax-like solutions, revealing a new theoretical bias in training dynamics.

Key results

2
Theorem 3.1 depth threshold

For L = 2 with K ≥ 6, or L ≥ 3 with K ≥ 4, deep neural collapse is not globally optimal in the max-margin landscape, though it remains a local optimum.

10
Figure 2 effective rank depth trend

In the K = 10 deep UFM experiment, the effective rank of the converged logit matrix decreases systematically as depth increases.

0
Theorem 4.5 width limit

In the joint limits ϵ → 0 and d → ∞, the initial logit derivative converges in probability to (L + 1)ϵ^{2L}(S ⊗ 1_n^T), showing large width biases early dynamics toward neural collapse.

What the paper found

The paper studies the deep unconstrained feature model under multiclass cross-entropy without explicit regularization, isolating the implicit bias of gradient flow in deep linear classifiers. Its main result is that depth itself induces a low-rank bias that competes with neural collapse: for L=2 and K≥6, or L≥3 and K≥4, the asymptotic max-margin landscape is no longer benign and deep neural collapse is not globally optimal, even though it remains a local optimum. The global optimum in the large-depth limit is a rank-2 softmax code, where class means lie uniformly on a circle, matching the d=2 bottleneck solution of Jiang et al. (2023). The mechanism is spectral: low-rank matrices propagate norm more efficiently through repeated products, so they achieve larger logits and smaller loss than the simplex ETF geometry associated with neural collapse. Dynamically, under Hadamard initialization the singular-value flow reduces to explicit ODEs showing a rich-get-richer effect, where larger singular modes grow faster near the origin, making neural collapse unstable for L>1 in the linearized regime. This instability grows with depth, while random Gaussian initialization at large width counteracts it by biasing early motion toward neural collapse through concentration of measure; the paper proves dZ/dt at time 0 converges in probability to (L+1)ϵ^{2L}(S⊗1_n^T). Experiments on deep UFMs and ResNet-20 heads trained on MNIST and CIFAR-10 confirm that increasing depth lowers effective rank, increasing width raises it, and lower-rank solutions tend to generalize worse.

Original abstract

Neural collapse (NC) describes the structured geometry that emerges in the features and weights of trained classifiers. Recent theory suggests NC can be suboptimal in deep architectures, attributing this to an explicit low-rank bias from L2 regularization. We study the deep unconstrained feature model (UFM)-equivalent to a deep linear network with orthogonal inputs-trained without regularization, to isolate how gradient descent and depth alone shape NC. We show that depth induces an implicit low-rank bias: low-rank matrices propagate norm more efficiently through successive multiplications, promoting low-rank alternatives to NC. These alternatives, we argue, correspond to softmax codes: max-margin solutions previously found in width-bottlenecked networks. Analyzing training dynamics under spectral initialization, we identify an early-time repulsion among singular values that drives low-rank emergence, and characterize how depth shrinks NC's basin of attraction. Finally, we show that some effects act in the opposite direction: for randomly initialized networks, increasing width biases training toward higher-rank solutions. Our results provide the first asymptotic and dynamic characterization of implicit bias in deep UFMs trained with unregularized multiclass cross-entropy.

Read the original paper

More in Neural Networks

Browse all 22 papers →
02Neural Network

Retrieving Individual Stems from Music Mixtures with Slot Embeddings

David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein

Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.

Read analysis
03Neural Network

The Linear Representation Hypothesis Needs a Group Action

Louie Hong Yao, Yuhao Li, Shengchao Liu

This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.

Read analysis