Distilled Reinforcement Learning for LLM Post-training
AuthorsChen Wang, Zhaochun Li, Jionghao Bai, Yining Zhang, Hexuan Deng, Ge Lan, Yue Wang
Distilled RL combines reinforcement learning with selective teacher guidance to help LLMs acquire new knowledge more effectively than standard RL or direct logit matching.
Key results
DAPO-17K was used for post-training.
Each training prompt generated a group of 8 student responses.
Average mathematical reasoning score in cross-family distillation.
Average mathematical reasoning score in within-family distillation.
Average Pass@1 decrease on Qwen3-4B when negative sample reset is removed.
What the paper found
Distilled Reinforcement Learning proposes replacing the unconditional reverse-KL imitation of on-policy distillation with teacher-guided policy optimization. For each student-generated trajectory, reverse importance sampling reweights tokens according to the Qwen3-8B-GRPO teacher, negative sample reset disables teacher weighting on responses with nonpositive advantage, and sequence-level geometric normalization preserves relative token preferences without globally scaling a response. Trained on DAPO-17K with rollout groups of 8, the method improves both within-family Qwen3 distillation and cross-family transfer to DeepSeek-R1-Distill-Qwen-1.5B, or DSQW-1.5B. On average Pass@1, DSQW-1.5B reaches 40.00, compared with 35.27 for OPD, while Qwen3-4B reaches 58.96, compared with 55.97 for OPD. The cross-family result is especially important because standard KL matching becomes unreliable when teacher and student have different reasoning distributions. Ablations show that removing negative sample reset lowers average Pass@1 by 8.81 points on Qwen3-4B, while removing geometric normalization also degrades performance. An entropy-controlled teacher experiment further indicates that Distilled RL can transfer distributional properties beyond the original task reward, offering a selective alternative to combining conventional reinforcement learning with a separate distillation loss.
Original abstract
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome supervision, resulting in difficult credit assignment and limited capability to acquire new knowledge. OPD, meanwhile, unconditionally matches teacher logits through KL divergence, which creates a dilemma: similar teachers provide little new knowledge, while substantially different teachers often yield ineffective guidance, largely restricting OPD to within-family distillation. We propose Distilled Reinforcement Learning (Distilled RL), which integrates teacher supervision into the RL objective to provide fine-grained guidance, selectively transfer new knowledge and avoid unconditional imitation. Distilled RL contains three components: reverse importance sampling with clipping, negative sample reset, and sequence-level geometric normalization. Through a concise and interpretable case study, we demonstrate that Distilled RL can effectively transfer previously unavailable knowledge from a teacher model to a student model. Extensive experiments across both within-family and cross-family distillation settings show that Distilled RL substantially outperforms standard RL and OPD in terms of both pass@1 and pass@k. Our code is available at https://github.com/597358816/Distilled-RL.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.