Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
AuthorsMostafa Elhoushi, Alex Pretko, Nolan Dey, Bin Claire Zhang, Gavia Gray, Gurpreet Gosal, Abdulrahman Mahmoud, Shane Bergsma, Joel Hestness
Resources
Layer dropout can make large language models cheaper to train and faster to run without meaningfully sacrificing accuracy.
Key results
Experiments spanning model and data scales.
Savings achieved with optimized layer dropout at large scale.
Validation loss for the 3.9B model with 20% FLOPs savings.
Draft & Verify inference speedup for the 8.2B model.
Draft & Verify inference speedup for the 3.9B model.
What the paper found
This paper revisits layer dropout, or stochastic depth, for modern large language models, challenging the practice seen in dense pretraining recipes associated with GPT-3, PaLM, and Meta’s Llama 3. Instead of dropping individual activations, the method skips entire transformer blocks per sequence, creating hardware-efficient structured sparsity. Across more than 2400 training experiments covering 271M to 8.2B-parameter models and datasets up to 160B tokens, the authors find that accuracy depends critically on configuration: dropout should increase with depth through Increasing Layer Distribution, then decrease linearly to zero over training through the Decreasing Time Schedule, with residual scaling by inverse layer density. This recipe can match or beat dense validation loss while reducing training computation by up to 25%, including a 3.9B model that achieves validation loss 1.745 with 20% FLOPs savings. The same pretraining induces depth elasticity, allowing early exit and intermediate-layer skipping without retraining; alternating dropout is strongest for non-contiguous layer skipping, while increasing dropout is better for early exit. For post-training Draft & Verify self-speculative decoding on XSUM, the 8.2B model reaches 1.55x speedup, and the 3.9B model reaches 1.54x, where the dense baseline achieves only 1.02x. Overall, layer dropout becomes a unified strategy for cheaper training and flexible inference, implemented on Cerebras CS-3 systems without architectural changes or auxiliary pretraining losses.
Original abstract
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformers. However, as models and datasets have scaled, dropout - particularly layer dropout - has largely disappeared from large language models (LLMs) pre-training recipes. While some prior work has reported that dropout can degrade accuracy, no comprehensive study has quantified, let alone mitigated, this effect. In this study, we show that layer dropout should be used in state-of-the-art LLM training, establishing best practices and scaling analysis for both training and post-training benefits. Concretely, with optimal layer distribution, time schedule, and optimizer hyperparameters, we observe that at the same training FLOPs layer dropout leads to lower loss. For a given number of training steps, LLMs can achieve lower or similar validation loss while saving upto 25% of training FLOPs. Moreover, layer dropout enables significant post-training optimizations, such as early exit, intermediate-layer skipping, and self-speculative decoding, yielding up to 1.5x inference speedup with negligible accuracy loss. Across more than 2400 training experiments, spanning models from 271M to 8.2B parameters and datasets up to 160B tokens, we demonstrate that these findings extend reliably to large-scale training regimes. All pre-training experiments were run on Cerebras CS-3 systems.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.