Generalization Dynamics of LM Pre-training
AuthorsJiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
AffiliationsUC Berkeley · Stanford University, Google DeepMind
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Key results
Accuracy on the answer-plus-one evaluation before a sharp mode switch.
Accuracy collapsed at the neighboring checkpoint.
Accuracy rebounded at the next reported checkpoint.
The selected 4.5T-token checkpoint scored higher than the 4.9T-token checkpoint at 29.8%.
The selected 4.5T-token checkpoint, compared with 21% for the 4.9T-token checkpoint.
What the paper found
This paper challenges the idea that language models steadily become more generalizable as pre-training progresses. Tracking OLMo3 and Apertus across six inexpensive behavioral evaluations, it finds repeated, abrupt “mode-hopping”: models switch between generalizing from examples and relying on shallow patterns. The probes cover flipped sentiment labels, repeated or sequential answers, truth versus plausibility, System 1 versus System 2 reasoning, coherent multi-hop personas, and fine-tuning generalization. One striking OLMo3-32B result swings from 81% accuracy at 2.17T tokens to 0% at 2.19T, then back to 81.7% at 2.21T. These shifts persist even after training far beyond Chinchilla-optimal budgets, are locally stable under a single gradient step, and are only mitigated—not eliminated—by checkpoint averaging. The paper proposes a capacity-allocation explanation: generalizable circuits compete with shallow ones, and each pre-training data window can influence which dominates. The findings connect to broader reasoning and alignment goals discussed around DeepSeek, Anthropic, and OpenAI. Practically, selecting an intermediate OLMo3-32B checkpoint at 4.5T tokens rather than 4.9T improved GPQA accuracy after math post-training to 36.3%, versus 29.8%, and improved robustness to prefilling attacks after general post-training to 53%, versus 21%. The study also shows that selecting pre-training data can steer and stabilize these dynamics; Claude-generated paraphrases help test whether mode-hopping generalizes across similar prompts.
Original abstract
People typically assume that LMs stably mature from pattern-matching parrots to generalizable intelligence during pre-training. We build a toy eval suite and show this mental model is wrong: throughout pre-training, LMs frequently and suddenly hop between parrot-like and intelligence-like computations. We call this mode-hopping. Across our suite, LMs suddenly latch onto memorized or in-context patterns instead of in-context learning, use System 1 instead of System 2 thinking, pick up what sounds true instead of what is true, fail at multi-hop persona QA, out-of-context reasoning, and emergent misalignment -- then just as suddenly revert and generalize. Mode-hopping is not explained by standard optimization dynamics: it is locally stable and cannot be fixed by checkpoint averaging. We instead think of it as a capacity allocation problem: in a capacity-bounded model, generalizable circuits must compete with the shallow ones learned early in training, and the data in each pre-training window may decide which circuits win. Our suite provides a new efficient lens on generalization. We demonstrate two concrete applications: (i) select intermediate pre-training checkpoints that strongly generalize reasoning and alignment, better than the final pre- or mid-training checkpoints, and (ii) select pre-training data that controls and stabilizes generalization dynamics.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
Shangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao, Xiaoyun Wang, Taylor W. Killian, Weitong Zhang
LSPD uses reinforcement-learning ideas to make LLM policy distillation more sample-efficient while preserving the diversity needed for stronger reasoning.