An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
AuthorsShangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao, Xiaoyun Wang, Taylor W. Killian, Weitong Zhang
Affiliations[
Resources
LSPD uses reinforcement-learning ideas to make LLM policy distillation more sample-efficient while preserving the diversity needed for stronger reasoning.
Key results
Average across 18 model–benchmark combinations.
Average across the six mathematical reasoning benchmarks and three Qwen3 teacher–student settings.
Replay-buffer variant’s average solution-coverage score.
Training steps needed to reach saturated performance, versus more than 40 for baselines.
Average across AMC23, AIME24, and AIME25, compared with 60.43 for EOPD.
What the paper found
This paper reframes on-policy distillation, or OPD, as KL-regularized reinforcement learning: the teacher-to-reference log-probability ratio becomes a token-level reward, and the student minimizes reverse KL against the teacher. Building on that identity, Least-Square Policy Distillation, or LSPD, uses robust quadratic matching between student and teacher log-probabilities, explicit entropy regularization for exploration, and optimistic reward estimation inspired by value-based RL. Unlike PPO-style OPD, LSPD supports multiple updates per rollout and historical trajectory reuse; its replay-buffer version, LSPD-RB, is fully off-policy. The theory gives logarithmic regret, O(log K), under bounded log-ratio rewards and suitable coverage. Across six benchmarks—MATH-500, Minerva, Olympiad-Bench, AMC23, AIME24, and AIME25—and three Qwen3 teacher–student scales, LSPD reaches an average Avg@16 score of 31.60 and Pass@16 score of 53.30, while LSPD-RB reaches Pass@16 of 54.86. LSPD-RB saturates after 10 training steps, compared with more than 40 for baselines, and its average Pass@64 is 61.94 versus 60.43 for EOPD, showing stronger solution diversity as sampling increases. The experiments use Qwen3 models and NVIDIA hardware, positioning LSPD as a rollout-efficient alternative to standard distillation for mathematical reasoning.
Original abstract
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that brings optimistic exploration and off-policy data reuse from value-based RL into policy distillation. LSPD preserves policy diversity through exploration while improving rollout efficiency by repeatedly learning from previously collected trajectories. Our theoretical analysis connects LSPD to optimistic value-based learning and shows that its idealized formulation achieves a sharp $\tilde{\mathcal O}(\log K)$ regret bound under online exploration. Empirically, LSPD consistently outperforms existing distillation baselines across six mathematical reasoning benchmarks and diverse teacher-student settings, with average gains of +1.59 points in Avg@16. Remarkably, through Pass@k evaluations up to k=64, we found that LSPD better preserves policy diversity by achieving stronger performance as k grows. Its fully off-policy variant achieves comparable performance to vanilla OPD using only the first 25% of rollout batches. Together, these results provide an RL perspective on OPD that offers both a principled interpretation and a practical route toward more effective and rollout-efficient language model distillation.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.