NTH

A Primer in Post-Training Reasoning Data: What We Know About How It Works

AuthorsYaoming Li, Guangxiang Zhao, Qilong Shi, Lin Sun, Xiangzheng Zhang, Tong Yang

June 8, 2026 2 min read
Watch on YouTube
The one-line take

A comprehensive primer that maps the emerging landscape of post-training reasoning data and explains what makes it work for modern large language models.

Key results

150+
public studies synthesized

The primer synthesizes over 150 key public studies and system reports on post-training reasoning data.

What the paper found

A Primer in Post-Training Reasoning Data: What We Know About How It Works is a survey from Peking University, Tsinghua University, and Qiyuan Tech that synthesizes more than 150 public studies to explain why post-training data, not just optimization, drives reasoning gains in systems like OpenAI’s o1-style models and DeepSeek’s DeepSeek-R1. The paper argues that reasoning data should be treated as verifier-bearing records, not simple prompt-response pairs, and organizes the field around a verifier-anchored taxonomy with three contracts—programmatic verification for math, code, and Lean proof corpora; environmental verification for tools, web, app, and software-agent trajectories; and judgment-required verification for safety, medical, and rubric-based tasks. Its main novelty is an attribution framework: quality depends on where supervision enters the trajectory, whether labels are outcome-only or step-level, how behavior is bounded, and what lineage is inherited from teachers, filters, and self-play anchors. The authors emphasize counterintuitive findings such as long chain-of-thought traces not implying faithful reasoning, “hard” examples being useful only relative to a base model and verifier, and cleaned successful agent transcripts erasing the failures and retries needed for credit assignment. They also separate scaling into reachable ceiling versus efficiency, showing that data uniqueness, verifier refresh, search topology, and inference-time budget can change measured capability without meaning the model itself improved.

Original abstract

Post-training has become a primary driver of recent progress in large reasoning models, and reasoning data are often the key variable determining whether this stage succeeds. Work on post-training reasoning data has grown rapidly, yet this literature remains scattered across dataset papers, reinforcement-learning recipes, reward-model studies, benchmarks, and frontier system reports. This paper is the first primer to synthesize over 150 key public studies and system reports on post-training reasoning data. We organize the field around four questions: what data objects exist, what makes them useful, how they are constructed, and how they scale. Together, this organization provides an attribution framework for future reasoning-data releases and post-training recipes.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis