NTH

The Devil is in the Condition Numbers: Why is GLU Better than non-GLU Structure?

AuthorsXingyu Lyu, Qianqian Xu, Zhiyong Yang, Peisong Wen, Qingming Huang

May 22, 2026 2 min read
Watch on YouTube
The one-line take

This paper argues that GLU layers work better because they make the training dynamics better conditioned, helping models learn faster rather than necessarily generalizing better.

Key results

O(n/d^2)
ReGLU NTK condition number

Under Gaussian inputs and LeCun initialization, the GLU/ReGLU NTK becomes more compact and the condition number improves to O(n/d^2), explaining faster convergence.

Θ(m)
Smallest eigenvalue scaling

The smallest NTK eigenvalue remains on the order of Θ(m) for both models, but the GLU case is larger than the non-GLU case, strengthening the lower spectral bound.

mostly above 0.05
Generalization gap p-values

Across MLP Mixer on CIFAR-10, ViT on Tiny ImageNet, and GPT-2 on FineWeb-Edu, permutation tests found no significant shift in generalization gap, with p-values mostly above 0.05.

MLP Mixer on CIFAR-10, ViT on Tiny ImageNet, GPT-2 on FineWeb-Edu
Benchmarks tested

These are the real-model experiments used to test whether GLU changes generalization gap; the paper reports similar gap behavior with and without GLU on all three benchmarks.

What the paper found

This paper explains why GLU and its variants, especially ReGLU and SwiGLU, outperform non-gated feedforward blocks by tracing the effect to NTK conditioning rather than to a smaller generalization gap. In a two-layer network analyzed in the neural tangent kernel regime, the authors derive an approximate factorization K̃ ≈ K ⊙ (XX⊤/d), showing that the GLU gate suppresses gradient correlation and makes the kernel spectrum more compact. Under Gaussian inputs and LeCun initialization, they prove that ReGLU changes the scaling of the largest NTK eigenvalue from Θ(mn) for ReLU to Θ(mn/d), while the smallest eigenvalue remains Θ(m) but becomes larger than the non-GLU case; this yields a smaller condition number and therefore faster gradient descent convergence. Their experiments confirm the spectral prediction on synthetic data and on real models: replacing FFN blocks in ViT and GPT-2 with GLU variants reduces NTK condition numbers across ReLU, GELU/GEGLU, and SiLU/SwiGLU settings. The paper also derives a loss-crossing phenomenon: non-GLU models can descend faster early in training, but GLU overtakes later because its larger λmin controls the slowest mode. In contrast, across MLP Mixer on CIFAR-10, ViT on Tiny ImageNet, and GPT-2 on FineWeb-Edu, the generalization gap Ltest−Ltrain remains statistically similar with and without GLU, with permutation-test p-values mostly above 0.05, indicating that GLU’s primary benefit is optimization speed, not gap reduction.

Original abstract

Gated Linear Units (GLU) and their variants are widely adopted in modern open-source large language model architectures and consistently outperform their non-gated counterparts, yet the underlying reasons for this advantage remain unclear. In this work, we study GLU by analyzing two-layer networks in the neural tangent kernel (NTK) regime. Our analysis reveals that the GLU structure reshapes the NTK spectrum, leading to a smaller condition number and a more compact eigenvalue distribution. Building on this finding, we further analyze the resulting training dynamics and show how the reshaped spectrum leads to faster convergence of GLU models, including a characteristic loss-crossing phenomenon observed between GLU and non-GLU models. Finally, we empirically observe that GLU has limited impact in reducing the generalization gap on various models, including ViT and GPT-2, suggesting that its primary benefit lies in accelerating optimization rather than reducing the generalization gap.

Read the original paper

More in Optimization

Browse all 36 papers →