NTH

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

AuthorsMingze Wang, Shuchen Zhu, Yuxin Fang, Binghui Li, Kai Shen, Shu Zhong

June 6, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows that the tiny scale vectors inside LLM normalization layers matter a lot for training, and proposes simple changes that make large language models learn better with almost no extra cost.

Key results

80640
scale vectors in Llama-1B

All scale vectors together contain 80,640 parameters in the 1.03B Llama model, only 7.84×10^-5 of total model size.

7.84e-5
parameter share of scale vectors

Scale vectors account for 7.84×10^-5 of the parameters in the 1.03B Llama model.

0.028
validation loss increase without scale vectors

Removing scale vectors from 0.12B Llama raises terminal validation loss by about 0.028 under the same peak learning rate.

0.015
validation loss increase after retuning

Even after retuning the peak learning rate for the model without scale vectors, terminal loss remains about 0.015 higher.

1.4
token efficiency gain

Without scale vectors, matching training quality requires about 1.4× more tokens.

1.04
runtime overhead

The unified scale-vector strategy adds only 1.04× wall-clock time on the Dense-1B benchmarked training run.

What the paper found

ByteDance Seed and Peking University present a systematic study of normalization scale vectors in modern LLMs, focusing on RMSNorm in Llama and Gemma3. The paper shows that these vectors are tiny—only 80,640 parameters in a 1.03B Llama model, or 7.84×10^-5 of total size—yet removing them raises validation loss by about 0.028 under matched learning rate and still by 0.015 after retuning, cutting token efficiency by roughly 1.4×. The key theoretical result is that in Pre-Norm Transformers, scale vectors do not add expressivity because their effect can be absorbed into the following linear map; instead, they speed training through a self-amplifying preconditioning dynamic on the effective weights. The authors then split normalization into Input-Norm and Output-Norm types and prove that weight decay should be applied only to Input-Norm scale vectors, while it is harmful for Output-Norm scale vectors that directly control expressivity. Building on this, they propose three lightweight upgrades: branch-specific scale vectors for attention and FFN branches, dual-sided or dual-normalized placement around linear maps, and magnitude-direction reparameterization in Euclidean or exponential form. On 0.12B to 2B dense Llama and LlamaMoE models, trained on industrial-scale token budgets with AdamW, Muon, and wsd schedules, the unified strategy consistently lowers terminal loss, often by more than 0.02 on MoE runs, with negligible runtime overhead of 1.04× and memory overhead of 1.01×.

Original abstract

Normalization layers in modern large language models (LLMs) consist of a deterministic normalization operation and a learnable scale vector. While the normalization operation has been extensively studied, the scale vector remains poorly understood despite its ubiquitous use. In this work, we present a systematic study of scale vectors in LLMs from the perspectives of expressivity, optimization, and architectural structure. First, we show empirically that although scale vectors constitute only a negligible fraction of model parameters, removing them substantially degrades LLM pre-training. Our theory further shows that, in Pre-Norm architectures, scale vectors do not increase expressivity; instead, they improve optimization through a self-amplifying preconditioning effect on subsequent linear mappings. Second, we investigate the role of weight decay for scale vectors. By distinguishing Input-Norm and Output-Norm layers, we theoretically show that weight decay is beneficial for the former but harmful for the latter, due to their distinct roles in optimization and expressivity. Third, motivated by this understanding, we propose three lightweight and complementary improvements to scale vectors: branch-specific heterogeneity, improved placement around linear mappings, and magnitude-direction reparameterization. Both theory and experiments show that each improvement yields consistent gains. Finally, we combine these improvements into a unified scale-vector strategy and evaluate it through extensive LLM pre-training experiments on dense and mixture-of-experts models ranging from 0.12B to 2B parameters, across multiple optimizers and learning rate schedules, under industrial-scale token budgets. The unified strategy consistently achieves lower terminal loss than well-tuned baselines and exhibits more favorable scaling behavior, while adding negligible parameter and computational overhead.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis