The Devil is in the Condition Numbers: Why is GLU Better than non-GLU Structure?
AuthorsXingyu Lyu, Qianqian Xu, Zhiyong Yang, Peisong Wen, Qingming Huang
Resources
This paper argues that GLU layers work better because they make the training dynamics better conditioned, helping models learn faster rather than necessarily generalizing better.
Key results
Under Gaussian inputs and LeCun initialization, the GLU/ReGLU NTK becomes more compact and the condition number improves to O(n/d^2), explaining faster convergence.
The smallest NTK eigenvalue remains on the order of Θ(m) for both models, but the GLU case is larger than the non-GLU case, strengthening the lower spectral bound.
Across MLP Mixer on CIFAR-10, ViT on Tiny ImageNet, and GPT-2 on FineWeb-Edu, permutation tests found no significant shift in generalization gap, with p-values mostly above 0.05.
These are the real-model experiments used to test whether GLU changes generalization gap; the paper reports similar gap behavior with and without GLU on all three benchmarks.
What the paper found
This paper explains why GLU and its variants, especially ReGLU and SwiGLU, outperform non-gated feedforward blocks by tracing the effect to NTK conditioning rather than to a smaller generalization gap. In a two-layer network analyzed in the neural tangent kernel regime, the authors derive an approximate factorization K̃ ≈ K ⊙ (XX⊤/d), showing that the GLU gate suppresses gradient correlation and makes the kernel spectrum more compact. Under Gaussian inputs and LeCun initialization, they prove that ReGLU changes the scaling of the largest NTK eigenvalue from Θ(mn) for ReLU to Θ(mn/d), while the smallest eigenvalue remains Θ(m) but becomes larger than the non-GLU case; this yields a smaller condition number and therefore faster gradient descent convergence. Their experiments confirm the spectral prediction on synthetic data and on real models: replacing FFN blocks in ViT and GPT-2 with GLU variants reduces NTK condition numbers across ReLU, GELU/GEGLU, and SiLU/SwiGLU settings. The paper also derives a loss-crossing phenomenon: non-GLU models can descend faster early in training, but GLU overtakes later because its larger λmin controls the slowest mode. In contrast, across MLP Mixer on CIFAR-10, ViT on Tiny ImageNet, and GPT-2 on FineWeb-Edu, the generalization gap Ltest−Ltrain remains statistically similar with and without GLU, with permutation-test p-values mostly above 0.05, indicating that GLU’s primary benefit is optimization speed, not gap reduction.
Original abstract
Gated Linear Units (GLU) and their variants are widely adopted in modern open-source large language model architectures and consistently outperform their non-gated counterparts, yet the underlying reasons for this advantage remain unclear. In this work, we study GLU by analyzing two-layer networks in the neural tangent kernel (NTK) regime. Our analysis reveals that the GLU structure reshapes the NTK spectrum, leading to a smaller condition number and a more compact eigenvalue distribution. Building on this finding, we further analyze the resulting training dynamics and show how the reshaped spectrum leads to faster convergence of GLU models, including a characteristic loss-crossing phenomenon observed between GLU and non-GLU models. Finally, we empirically observe that GLU has limited impact in reducing the generalization gap on various models, including ViT and GPT-2, suggesting that its primary benefit lies in accelerating optimization rather than reducing the generalization gap.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.