NTH

Understanding Reasoning from Pretraining to Post-Training

AuthorsJingyan Shen, Ang Li, Salman Rahman, Yifan Sun, Micah Goldblum, Matus Telgarsky, Pavel Izmailov

July 21, 2026 3 min read
Watch on YouTube
The one-line take

Using chess and math as controlled laboratories, the paper shows how pretraining sets the stage for—and shapes what RL can achieve in—LLM reasoning.

Key results

54B
Chess pretraining corpus

Lichess human-game tokens used for autoregressive pretraining.

36
Pretraining–RL sweep

Number of pretraining and reinforcement-learning combinations evaluated.

0.84
Token–RL-slope correlation

Spearman correlation between log pretraining tokens and the local RL improvement slope.

28%
RL share at high-compute frontier

Predicted optimal RL fraction at the 680M model frontier, up from approximately 20% at 50M.

200B
Math pretraining scale

Largest pretraining-token checkpoint evaluated for the 1B-parameter OLMo-2 transfer experiment.

What the paper found

A team led by New York University, with researchers from Modal Labs, UCLA, UIUC, and Columbia, studies how reasoning develops from pretraining through supervised fine-tuning and reinforcement learning. Using dense Qwen3-style models ranging from 5M to 1B parameters, they build a controlled chess pipeline: autoregressive pretraining on a 54B-token Lichess corpus, synthetic tree-structured reasoning traces for SFT, and GRPO with binary, verifiable puzzle rewards. Across 36 pretraining–RL combinations, the researchers derive a joint scaling law in which lower pretraining loss predicts higher post-RL pass@1, while the RL improvement slope grows mainly with log pretraining tokens and only weakly with model size; the token–slope relationship reaches a Spearman correlation of 0.84. Compute-optimal allocations shift toward RL as budgets grow, with the predicted RL share rising from approximately 20% to 28% between the 50M and 680M model frontiers, while Chinchilla-like token allocation remains broadly stable. Mechanistically, RL is not simple uniform sharpening: on easy puzzles it amplifies correct moves already favored by SFT, but on hard puzzles it can discover correct tail moves while also reinforcing wrong modes, explaining stronger pass@1 gains without consistent pass@16 gains. A transfer experiment with a 1B-parameter OLMo-2 model trained on math data from 10B to 200B tokens reproduces the same pattern, suggesting the pretraining–RL interface extends beyond chess and beyond systems such as DeepSeek-R1, while relying on familiar components including NVIDIA H200 hardware, Qwen3, OLMo-2, GSM8K, and MATH.

Original abstract

Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM setting: pretraining corpora are vast and uncontrolled, making it hard to attribute behaviors to pretraining versus RL, and systematic compute sweeps across both stages are prohibitively expensive. To address these challenges, we use chess as a controlled testbed for studying reasoning across the full pretraining-to-post-training pipeline. We follow the standard LLM training pipeline by pretraining language models from 5M to 1B parameters on human chess games, supervised fine-tuning on synthetic reasoning traces, and running RL on chess puzzles with verifiable rewards. Using this framework, we find that the post-RL performance at given RL compute level is well-predicted from the pretraining loss, and slope of the RL reward curves improves approximately linearly with the pretraining tokens. Beyond scaling, we find that RL does not simply sharpen the SFT policy: on easy puzzles it amplifies correct moves the SFT policy already preferred, while on hard puzzles it surfaces correct moves that were nearly absent under SFT. We further test whether our findings transfer beyond chess by training a 1B language model on math-domain text, where the same predictive pattern emerges: longer-pretrained checkpoints reach higher post-RL performance and improve faster under RL. In sum, we provide a quantitative account of the pretraining-to-RL interface and a controlled testbed for studying the science of reasoning across the full pretraining-to-post-training pipeline.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →