A Primer in Post-Training Reasoning Data: What We Know About How It Works
AuthorsYaoming Li, Guangxiang Zhao, Qilong Shi, Lin Sun, Xiangzheng Zhang, Tong Yang
Resources
A comprehensive primer that maps the emerging landscape of post-training reasoning data and explains what makes it work for modern large language models.
Key results
The primer synthesizes over 150 key public studies and system reports on post-training reasoning data.
What the paper found
A Primer in Post-Training Reasoning Data: What We Know About How It Works is a survey from Peking University, Tsinghua University, and Qiyuan Tech that synthesizes more than 150 public studies to explain why post-training data, not just optimization, drives reasoning gains in systems like OpenAI’s o1-style models and DeepSeek’s DeepSeek-R1. The paper argues that reasoning data should be treated as verifier-bearing records, not simple prompt-response pairs, and organizes the field around a verifier-anchored taxonomy with three contracts—programmatic verification for math, code, and Lean proof corpora; environmental verification for tools, web, app, and software-agent trajectories; and judgment-required verification for safety, medical, and rubric-based tasks. Its main novelty is an attribution framework: quality depends on where supervision enters the trajectory, whether labels are outcome-only or step-level, how behavior is bounded, and what lineage is inherited from teachers, filters, and self-play anchors. The authors emphasize counterintuitive findings such as long chain-of-thought traces not implying faithful reasoning, “hard” examples being useful only relative to a base model and verifier, and cleaned successful agent transcripts erasing the failures and retries needed for credit assignment. They also separate scaling into reachable ceiling versus efficiency, showing that data uniqueness, verifier refresh, search topology, and inference-time budget can change measured capability without meaning the model itself improved.
Original abstract
Post-training has become a primary driver of recent progress in large reasoning models, and reasoning data are often the key variable determining whether this stage succeeds. Work on post-training reasoning data has grown rapidly, yet this literature remains scattered across dataset papers, reinforcement-learning recipes, reward-model studies, benchmarks, and frontier system reports. This paper is the first primer to synthesize over 150 key public studies and system reports on post-training reasoning data. We organize the field around four questions: what data objects exist, what makes them useful, how they are constructed, and how they scale. Together, this organization provides an attribution framework for future reasoning-data releases and post-training recipes.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.