NTH

Generalization Dynamics of LM Pre-training

AuthorsJiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

AffiliationsUC Berkeley · Stanford University, Google DeepMind

October 9, 2026 2 min read
Watch on YouTube
The one-line take

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Key results

81%
OLMo3-32B successive-answer accuracy at 2.17T tokens

Accuracy on the answer-plus-one evaluation before a sharp mode switch.

0%
OLMo3-32B successive-answer accuracy at 2.19T tokens

Accuracy collapsed at the neighboring checkpoint.

81.7%
OLMo3-32B successive-answer accuracy at 2.21T tokens

Accuracy rebounded at the next reported checkpoint.

36.3%
GPQA accuracy after math post-training

The selected 4.5T-token checkpoint scored higher than the 4.9T-token checkpoint at 29.8%.

53%
Robustness to prefilling attacks after general post-training

The selected 4.5T-token checkpoint, compared with 21% for the 4.9T-token checkpoint.

What the paper found

This paper challenges the idea that language models steadily become more generalizable as pre-training progresses. Tracking OLMo3 and Apertus across six inexpensive behavioral evaluations, it finds repeated, abrupt “mode-hopping”: models switch between generalizing from examples and relying on shallow patterns. The probes cover flipped sentiment labels, repeated or sequential answers, truth versus plausibility, System 1 versus System 2 reasoning, coherent multi-hop personas, and fine-tuning generalization. One striking OLMo3-32B result swings from 81% accuracy at 2.17T tokens to 0% at 2.19T, then back to 81.7% at 2.21T. These shifts persist even after training far beyond Chinchilla-optimal budgets, are locally stable under a single gradient step, and are only mitigated—not eliminated—by checkpoint averaging. The paper proposes a capacity-allocation explanation: generalizable circuits compete with shallow ones, and each pre-training data window can influence which dominates. The findings connect to broader reasoning and alignment goals discussed around DeepSeek, Anthropic, and OpenAI. Practically, selecting an intermediate OLMo3-32B checkpoint at 4.5T tokens rather than 4.9T improved GPQA accuracy after math post-training to 36.3%, versus 29.8%, and improved robustness to prefilling attacks after general post-training to 53%, versus 21%. The study also shows that selecting pre-training data can steer and stabilize these dynamics; Claude-generated paraphrases help test whether mode-hopping generalizes across similar prompts.

Original abstract

People typically assume that LMs stably mature from pattern-matching parrots to generalizable intelligence during pre-training. We build a toy eval suite and show this mental model is wrong: throughout pre-training, LMs frequently and suddenly hop between parrot-like and intelligence-like computations. We call this mode-hopping. Across our suite, LMs suddenly latch onto memorized or in-context patterns instead of in-context learning, use System 1 instead of System 2 thinking, pick up what sounds true instead of what is true, fail at multi-hop persona QA, out-of-context reasoning, and emergent misalignment -- then just as suddenly revert and generalize. Mode-hopping is not explained by standard optimization dynamics: it is locally stable and cannot be fixed by checkpoint averaging. We instead think of it as a capacity allocation problem: in a capacity-bounded model, generalizable circuits must compete with the shallow ones learned early in training, and the data in each pre-training window may decide which circuits win. Our suite provides a new efficient lens on generalization. We demonstrate two concrete applications: (i) select intermediate pre-training checkpoints that strongly generalize reasoning and alignment, better than the final pre- or mid-training checkpoints, and (ii) select pre-training data that controls and stabilizes generalization dynamics.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis