NTH

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

AuthorsJacqueline He, Howard Yen, Shuyue Stella Li, Margaret Li, Hanqing Zeng, Yinglong Xia, Benyu Zhang, Zhuokai Zhao, Qiang Zhang, Pang Wei Koh, Luke Zettlemoyer, Wen-tau Yih

September 9, 2026 2 min read
Watch on YouTube
The one-line take

A new distillation strategy uses teacher confidence to help language models gain reasoning skills without sacrificing the factual knowledge they learn during mid-training.

Key results

60B
Mid-training token budget

Tokens used to continue training the 1B OLMo-2 student from a 4T-token checkpoint.

20%
Switch routing threshold

Lowest-entropy tokens routed to reverse-KL distillation.

26.1
NTP reasoning baseline

Mid-training macro-average reasoning performance under standard next-token prediction.

44.7
Switch Distillation reasoning

Mid-training reasoning macro-average with the OLMo-2 7B Instruct teacher.

29.3
Switch Distillation factual recall

Mid-training factual-recall macro-average, compared with 30.3% for NTP.

50.6
Post-training reasoning

Final reasoning macro-average after post-training with the 7B teacher.

What the paper found

This paper examines logit-based knowledge distillation across language-model training stages using a 1B OLMo-2 student, OLMo-2 Instruct teachers at 1B, 7B, and 13B, and the Dolmino Mix 1124 corpus. Standard forward-KL distillation improves both reasoning and factual recall during pre-training, but during mid-training it creates a reasoning–recall tradeoff: on a 60B-token continuation from a 4T-token checkpoint, reasoning improves while factual acquisition slows. The mechanism is teacher-entropy asymmetry: instruction and mathematics tokens receive confident, useful supervision, whereas unresolved factual tokens produce diffuse teacher distributions, weakening the ground-truth learning signal to roughly 0.5× the next-token-prediction signal at the highest entropy. The proposed Switch Distillation routes the lowest-entropy 20% of tokens to reverse-KL distillation and trains the remainder with cross-entropy, using no additional parameters or model passes. With a 7B teacher, it raises mid-training reasoning from 26.1% for standard next-token prediction to 44.7%, while factual recall reaches 29.3% versus 30.3% for the baseline; it also reaches 49.3% on Knowledge and Commonsense. After supervised fine-tuning, direct preference optimization, and reinforcement learning with verifiable rewards, reasoning rises to 50.6%, and the factual-recall gap closes while the reasoning advantage persists. The same qualitative pattern appears in SmolLM2, and entropy separation generalizes across Qwen3, Gemma-3, and Granite 3.3 teachers.

Original abstract

Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits are consistent across training stages remains unclear. Through controlled experiments, we find that forward Kullback-Leibler (KL) distillation--the standard KD formulation--with post-trained teachers behaves fundamentally differently during mid-training, an intermediate phase of self-supervised learning on curated corpora. Surprisingly, while forward KD simultaneously improves reasoning and factual recall during pre-training relative to standard next-token prediction (NTP), it instead slows factual recall acquisition during mid-training despite continued reasoning gains. We trace this stage dependence to an asymmetry in teacher confidence across data domains and the student's evolving knowledge state: teachers are more confident on procedural than knowledge-intensive data, while students acquire low-entropy factual knowledge earlier in training. To mitigate this imbalance, we propose Switch Distillation, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy. Switch Distillation consistently outperforms existing distillation objectives across teacher sizes. Relative to standard NTP, it achieves 1.61-1.71x the reasoning performance and 1.13-1.19x the knowledge and commonsense performance while preserving 96.7-96.8% of factual recall. Crucially, these benefits persist after post-training: Switch Distillation closes the factual recall gap while maintaining 1.25-1.32x and 1.13-1.20x gains in reasoning and knowledge and commonsense, respectively.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis