NTH

An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning

AuthorsShangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao, Xiaoyun Wang, Taylor W. Killian, Weitong Zhang

Affiliations[

October 8, 2026 2 min read
Watch on YouTube
The one-line take

LSPD uses reinforcement-learning ideas to make LLM policy distillation more sample-efficient while preserving the diversity needed for stronger reasoning.

Key results

31.60
LSPD average Avg@16

Average across 18 model–benchmark combinations.

53.30
LSPD average Pass@16

Average across the six mathematical reasoning benchmarks and three Qwen3 teacher–student settings.

54.86
LSPD-RB average Pass@16

Replay-buffer variant’s average solution-coverage score.

10
LSPD-RB convergence steps

Training steps needed to reach saturated performance, versus more than 40 for baselines.

61.94
LSPD average Pass@64

Average across AMC23, AIME24, and AIME25, compared with 60.43 for EOPD.

What the paper found

This paper reframes on-policy distillation, or OPD, as KL-regularized reinforcement learning: the teacher-to-reference log-probability ratio becomes a token-level reward, and the student minimizes reverse KL against the teacher. Building on that identity, Least-Square Policy Distillation, or LSPD, uses robust quadratic matching between student and teacher log-probabilities, explicit entropy regularization for exploration, and optimistic reward estimation inspired by value-based RL. Unlike PPO-style OPD, LSPD supports multiple updates per rollout and historical trajectory reuse; its replay-buffer version, LSPD-RB, is fully off-policy. The theory gives logarithmic regret, O(log K), under bounded log-ratio rewards and suitable coverage. Across six benchmarks—MATH-500, Minerva, Olympiad-Bench, AMC23, AIME24, and AIME25—and three Qwen3 teacher–student scales, LSPD reaches an average Avg@16 score of 31.60 and Pass@16 score of 53.30, while LSPD-RB reaches Pass@16 of 54.86. LSPD-RB saturates after 10 training steps, compared with more than 40 for baselines, and its average Pass@64 is 61.94 versus 60.43 for EOPD, showing stronger solution diversity as sampling increases. The experiments use Qwen3 models and NVIDIA hardware, positioning LSPD as a rollout-efficient alternative to standard distillation for mathematical reasoning.

Original abstract

We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that brings optimistic exploration and off-policy data reuse from value-based RL into policy distillation. LSPD preserves policy diversity through exploration while improving rollout efficiency by repeatedly learning from previously collected trajectories. Our theoretical analysis connects LSPD to optimistic value-based learning and shows that its idealized formulation achieves a sharp $\tilde{\mathcal O}(\log K)$ regret bound under online exploration. Empirically, LSPD consistently outperforms existing distillation baselines across six mathematical reasoning benchmarks and diverse teacher-student settings, with average gains of +1.59 points in Avg@16. Remarkably, through Pass@k evaluations up to k=64, we found that LSPD better preserves policy diversity by achieving stronger performance as k grows. Its fully off-policy variant achieves comparable performance to vanilla OPD using only the first 25% of rollout batches. Together, these results provide an RL perspective on OPD that offers both a principled interpretation and a practical route toward more effective and rollout-efficient language model distillation.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis