Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
AuthorsJacqueline He, Howard Yen, Shuyue Stella Li, Margaret Li, Hanqing Zeng, Yinglong Xia, Benyu Zhang, Zhuokai Zhao, Qiang Zhang, Pang Wei Koh, Luke Zettlemoyer, Wen-tau Yih
Resources
A new distillation strategy uses teacher confidence to help language models gain reasoning skills without sacrificing the factual knowledge they learn during mid-training.
Key results
Tokens used to continue training the 1B OLMo-2 student from a 4T-token checkpoint.
Lowest-entropy tokens routed to reverse-KL distillation.
Mid-training macro-average reasoning performance under standard next-token prediction.
Mid-training reasoning macro-average with the OLMo-2 7B Instruct teacher.
Mid-training factual-recall macro-average, compared with 30.3% for NTP.
Final reasoning macro-average after post-training with the 7B teacher.
What the paper found
This paper examines logit-based knowledge distillation across language-model training stages using a 1B OLMo-2 student, OLMo-2 Instruct teachers at 1B, 7B, and 13B, and the Dolmino Mix 1124 corpus. Standard forward-KL distillation improves both reasoning and factual recall during pre-training, but during mid-training it creates a reasoning–recall tradeoff: on a 60B-token continuation from a 4T-token checkpoint, reasoning improves while factual acquisition slows. The mechanism is teacher-entropy asymmetry: instruction and mathematics tokens receive confident, useful supervision, whereas unresolved factual tokens produce diffuse teacher distributions, weakening the ground-truth learning signal to roughly 0.5× the next-token-prediction signal at the highest entropy. The proposed Switch Distillation routes the lowest-entropy 20% of tokens to reverse-KL distillation and trains the remainder with cross-entropy, using no additional parameters or model passes. With a 7B teacher, it raises mid-training reasoning from 26.1% for standard next-token prediction to 44.7%, while factual recall reaches 29.3% versus 30.3% for the baseline; it also reaches 49.3% on Knowledge and Commonsense. After supervised fine-tuning, direct preference optimization, and reinforcement learning with verifiable rewards, reasoning rises to 50.6%, and the factual-recall gap closes while the reasoning advantage persists. The same qualitative pattern appears in SmolLM2, and entropy separation generalizes across Qwen3, Gemma-3, and Granite 3.3 teachers.
Original abstract
Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits are consistent across training stages remains unclear. Through controlled experiments, we find that forward Kullback-Leibler (KL) distillation--the standard KD formulation--with post-trained teachers behaves fundamentally differently during mid-training, an intermediate phase of self-supervised learning on curated corpora. Surprisingly, while forward KD simultaneously improves reasoning and factual recall during pre-training relative to standard next-token prediction (NTP), it instead slows factual recall acquisition during mid-training despite continued reasoning gains. We trace this stage dependence to an asymmetry in teacher confidence across data domains and the student's evolving knowledge state: teachers are more confident on procedural than knowledge-intensive data, while students acquire low-entropy factual knowledge earlier in training. To mitigate this imbalance, we propose Switch Distillation, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy. Switch Distillation consistently outperforms existing distillation objectives across teacher sizes. Relative to standard NTP, it achieves 1.61-1.71x the reasoning performance and 1.13-1.19x the knowledge and commonsense performance while preserving 96.7-96.8% of factual recall. Crucially, these benefits persist after post-training: Switch Distillation closes the factual recall gap while maintaining 1.25-1.32x and 1.13-1.20x gains in reasoning and knowledge and commonsense, respectively.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.