NTH

Distilled Reinforcement Learning for LLM Post-training

AuthorsChen Wang, Zhaochun Li, Jionghao Bai, Yining Zhang, Hexuan Deng, Ge Lan, Yue Wang

July 24, 2026 2 min read
Watch on YouTube
The one-line take

Distilled RL combines reinforcement learning with selective teacher guidance to help LLMs acquire new knowledge more effectively than standard RL or direct logit matching.

Key results

17K
Training dataset scale

DAPO-17K was used for post-training.

8
Rollout group size

Each training prompt generated a group of 8 student responses.

40.00
DSQW-1.5B Distilled RL average Pass@1

Average mathematical reasoning score in cross-family distillation.

58.96
Qwen3-4B Distilled RL average Pass@1

Average mathematical reasoning score in within-family distillation.

8.81
Negative-reset ablation drop

Average Pass@1 decrease on Qwen3-4B when negative sample reset is removed.

What the paper found

Distilled Reinforcement Learning proposes replacing the unconditional reverse-KL imitation of on-policy distillation with teacher-guided policy optimization. For each student-generated trajectory, reverse importance sampling reweights tokens according to the Qwen3-8B-GRPO teacher, negative sample reset disables teacher weighting on responses with nonpositive advantage, and sequence-level geometric normalization preserves relative token preferences without globally scaling a response. Trained on DAPO-17K with rollout groups of 8, the method improves both within-family Qwen3 distillation and cross-family transfer to DeepSeek-R1-Distill-Qwen-1.5B, or DSQW-1.5B. On average Pass@1, DSQW-1.5B reaches 40.00, compared with 35.27 for OPD, while Qwen3-4B reaches 58.96, compared with 55.97 for OPD. The cross-family result is especially important because standard KL matching becomes unreliable when teacher and student have different reasoning distributions. Ablations show that removing negative sample reset lowers average Pass@1 by 8.81 points on Qwen3-4B, while removing geometric normalization also degrades performance. An entropy-controlled teacher experiment further indicates that Distilled RL can transfer distributional properties beyond the original task reward, offering a selective alternative to combining conventional reinforcement learning with a separate distillation loss.

Original abstract

Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome supervision, resulting in difficult credit assignment and limited capability to acquire new knowledge. OPD, meanwhile, unconditionally matches teacher logits through KL divergence, which creates a dilemma: similar teachers provide little new knowledge, while substantially different teachers often yield ineffective guidance, largely restricting OPD to within-family distillation. We propose Distilled Reinforcement Learning (Distilled RL), which integrates teacher supervision into the RL objective to provide fine-grained guidance, selectively transfer new knowledge and avoid unconditional imitation. Distilled RL contains three components: reverse importance sampling with clipping, negative sample reset, and sequence-level geometric normalization. Through a concise and interpretable case study, we demonstrate that Distilled RL can effectively transfer previously unavailable knowledge from a teacher model to a student model. Extensive experiments across both within-family and cross-family distillation settings show that Distilled RL substantially outperforms standard RL and OPD in terms of both pass@1 and pass@k. Our code is available at https://github.com/597358816/Distilled-RL.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis