NTH

dOPSD: On-Policy Self-Distillation for Diffusion Language Models

AuthorsPhuong Tuan Dat, Qi Li, Xinchao Wang

July 9, 2026 2 min read
Watch on YouTube
The one-line take

dOPSD teaches diffusion language models using their own unfolding predictions, boosting reasoning without needing ground-truth hints.

Key results

81.41
Dream GSM8K base

Base accuracy on Dream-7B-Instruct

83.04
Dream GSM8K dOPSD

dOPSD accuracy on Dream-7B-Instruct

38.97
Dream MATH500 base

Base accuracy on Dream-7B-Instruct

42.20
Dream MATH500 dOPSD

dOPSD accuracy on Dream-7B-Instruct

52.54
Dream HumanEval base

Base pass@1 on Dream-7B-Instruct

56.71
Dream HumanEval dOPSD

dOPSD pass@1 on Dream-7B-Instruct

What the paper found

dOPSD, from National University of Singapore, adapts on-policy self-distillation to diffusion language models by replacing the external privileged reference used in OPSD with a privilege drawn from the model’s own denoising trajectory: the student is trained on genuine intermediate decode states, while the same model acting later in the trajectory serves as a teacher that averages predictions over future masked steps and distills them with a token-level generalized Jensen–Shannon objective. This matters because naive OPSD collapses in diffusion settings: on Dream-7B-Instruct and LLaDA-8B-Instruct, the answer-only OPSD variant barely moves from the base model, and the full-solution variant sharply degrades, including a 14.2-point GSM8K drop on Dream and a 13.4-point drop on LLaDA, confirming the paper’s claim that instance-specific privileged information becomes a weak PI-free consensus. By contrast, dOPSD is the only method that improves every benchmark in the study, raising Dream-7B-Instruct from 81.41 to 83.04 on GSM8K and from 38.97 to 42.20 on MATH500, while also improving out-of-distribution code generation from 52.54 to 56.71 on HumanEval and from 57.43 to 58.49 on MBPP. On LLaDA-8B-Instruct it lifts GSM8K from 71.23 to 72.87 and MATH500 from 31.24 to 36.00, again transferring to HumanEval and MBPP. Ablations show that a forward-KL target, a full future trajectory window, and a mask threshold of 0.5 are the strongest settings, and even without rollout verification dOPSD still beats the base model, indicating the main gain comes from trajectory-derived, on-policy privileged supervision rather than external labels.

Original abstract

Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel alternative to autoregressive models, but eliciting strong reasoning through post-training remains difficult: supervised fine-tuning is off-policy and suffers from exposure bias, while reinforcement learning gives only sparse, sequence-level rewards and is hard to apply without tractable sequence likelihoods. On-policy self-distillation (OPSD) offers a promising alternative, using one model as both student and teacher to provide dense, token-level, on-policy supervision, but its effectiveness hinges on giving the teacher privileged information (PI) - typically an instance-specific ground-truth reference unavailable at inference - so the student ends up distilling a weak PI-free consensus policy that yields little improvement on dLLM reasoning. We introduce dOPSD, which instead derives the teacher's privilege directly from the student's own denoising trajectory, evaluating masked positions using later, more-decoded steps of that same trajectory rather than an external label, so the teacher's advantage emerges from the model's own decoding process; on Dream and LLaDA, dOPSD improves both in-domain math reasoning and out-of-domain code generation, outperforming supervised and on-policy baselines.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis