dOPSD: On-Policy Self-Distillation for Diffusion Language Models
AuthorsPhuong Tuan Dat, Qi Li, Xinchao Wang
Resources
dOPSD teaches diffusion language models using their own unfolding predictions, boosting reasoning without needing ground-truth hints.
Key results
Base accuracy on Dream-7B-Instruct
dOPSD accuracy on Dream-7B-Instruct
Base accuracy on Dream-7B-Instruct
dOPSD accuracy on Dream-7B-Instruct
Base pass@1 on Dream-7B-Instruct
dOPSD pass@1 on Dream-7B-Instruct
What the paper found
dOPSD, from National University of Singapore, adapts on-policy self-distillation to diffusion language models by replacing the external privileged reference used in OPSD with a privilege drawn from the model’s own denoising trajectory: the student is trained on genuine intermediate decode states, while the same model acting later in the trajectory serves as a teacher that averages predictions over future masked steps and distills them with a token-level generalized Jensen–Shannon objective. This matters because naive OPSD collapses in diffusion settings: on Dream-7B-Instruct and LLaDA-8B-Instruct, the answer-only OPSD variant barely moves from the base model, and the full-solution variant sharply degrades, including a 14.2-point GSM8K drop on Dream and a 13.4-point drop on LLaDA, confirming the paper’s claim that instance-specific privileged information becomes a weak PI-free consensus. By contrast, dOPSD is the only method that improves every benchmark in the study, raising Dream-7B-Instruct from 81.41 to 83.04 on GSM8K and from 38.97 to 42.20 on MATH500, while also improving out-of-distribution code generation from 52.54 to 56.71 on HumanEval and from 57.43 to 58.49 on MBPP. On LLaDA-8B-Instruct it lifts GSM8K from 71.23 to 72.87 and MATH500 from 31.24 to 36.00, again transferring to HumanEval and MBPP. Ablations show that a forward-KL target, a full future trajectory window, and a mask threshold of 0.5 are the strongest settings, and even without rollout verification dOPSD still beats the base model, indicating the main gain comes from trajectory-derived, on-policy privileged supervision rather than external labels.
Original abstract
Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel alternative to autoregressive models, but eliciting strong reasoning through post-training remains difficult: supervised fine-tuning is off-policy and suffers from exposure bias, while reinforcement learning gives only sparse, sequence-level rewards and is hard to apply without tractable sequence likelihoods. On-policy self-distillation (OPSD) offers a promising alternative, using one model as both student and teacher to provide dense, token-level, on-policy supervision, but its effectiveness hinges on giving the teacher privileged information (PI) - typically an instance-specific ground-truth reference unavailable at inference - so the student ends up distilling a weak PI-free consensus policy that yields little improvement on dLLM reasoning. We introduce dOPSD, which instead derives the teacher's privilege directly from the student's own denoising trajectory, evaluating masked positions using later, more-decoded steps of that same trajectory rather than an external label, so the teacher's advantage emerges from the model's own decoding process; on Dream and LLaDA, dOPSD improves both in-domain math reasoning and out-of-domain code generation, outperforming supervised and on-policy baselines.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.