Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models
AuthorsMingze Wang, Shuchen Zhu, Yuxin Fang, Binghui Li, Kai Shen, Shu Zhong
Resources
This paper shows that the tiny scale vectors inside LLM normalization layers matter a lot for training, and proposes simple changes that make large language models learn better with almost no extra cost.
Key results
All scale vectors together contain 80,640 parameters in the 1.03B Llama model, only 7.84×10^-5 of total model size.
Scale vectors account for 7.84×10^-5 of the parameters in the 1.03B Llama model.
Removing scale vectors from 0.12B Llama raises terminal validation loss by about 0.028 under the same peak learning rate.
Even after retuning the peak learning rate for the model without scale vectors, terminal loss remains about 0.015 higher.
Without scale vectors, matching training quality requires about 1.4× more tokens.
The unified scale-vector strategy adds only 1.04× wall-clock time on the Dense-1B benchmarked training run.
What the paper found
ByteDance Seed and Peking University present a systematic study of normalization scale vectors in modern LLMs, focusing on RMSNorm in Llama and Gemma3. The paper shows that these vectors are tiny—only 80,640 parameters in a 1.03B Llama model, or 7.84×10^-5 of total size—yet removing them raises validation loss by about 0.028 under matched learning rate and still by 0.015 after retuning, cutting token efficiency by roughly 1.4×. The key theoretical result is that in Pre-Norm Transformers, scale vectors do not add expressivity because their effect can be absorbed into the following linear map; instead, they speed training through a self-amplifying preconditioning dynamic on the effective weights. The authors then split normalization into Input-Norm and Output-Norm types and prove that weight decay should be applied only to Input-Norm scale vectors, while it is harmful for Output-Norm scale vectors that directly control expressivity. Building on this, they propose three lightweight upgrades: branch-specific scale vectors for attention and FFN branches, dual-sided or dual-normalized placement around linear maps, and magnitude-direction reparameterization in Euclidean or exponential form. On 0.12B to 2B dense Llama and LlamaMoE models, trained on industrial-scale token budgets with AdamW, Muon, and wsd schedules, the unified strategy consistently lowers terminal loss, often by more than 0.02 on MoE runs, with negligible runtime overhead of 1.04× and memory overhead of 1.01×.
Original abstract
Normalization layers in modern large language models (LLMs) consist of a deterministic normalization operation and a learnable scale vector. While the normalization operation has been extensively studied, the scale vector remains poorly understood despite its ubiquitous use. In this work, we present a systematic study of scale vectors in LLMs from the perspectives of expressivity, optimization, and architectural structure. First, we show empirically that although scale vectors constitute only a negligible fraction of model parameters, removing them substantially degrades LLM pre-training. Our theory further shows that, in Pre-Norm architectures, scale vectors do not increase expressivity; instead, they improve optimization through a self-amplifying preconditioning effect on subsequent linear mappings. Second, we investigate the role of weight decay for scale vectors. By distinguishing Input-Norm and Output-Norm layers, we theoretically show that weight decay is beneficial for the former but harmful for the latter, due to their distinct roles in optimization and expressivity. Third, motivated by this understanding, we propose three lightweight and complementary improvements to scale vectors: branch-specific heterogeneity, improved placement around linear mappings, and magnitude-direction reparameterization. Both theory and experiments show that each improvement yields consistent gains. Finally, we combine these improvements into a unified scale-vector strategy and evaluate it through extensive LLM pre-training experiments on dense and mixture-of-experts models ranging from 0.12B to 2B parameters, across multiple optimizers and learning rate schedules, under industrial-scale token budgets. The unified strategy consistently achieves lower terminal loss than well-tuned baselines and exhibits more favorable scaling behavior, while adding negligible parameter and computational overhead.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.