Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy Distillation
AuthorsXingyu Su, Jacob Helwig, Shubham Parashar, Atharv Chagi, Lakshmi Jotsna, Degui Zhi, James Caverlee, Dileep Kalathil, Shuiwang Ji
Resources
This paper shows how to turn an autoregressive language model into a diffusion-style language model much more efficiently by training it on its own generated trajectories while distilling knowledge from the original model.
Key results
Training data used for the main 8B model
Training compute in 10^18 FLOPs for the main 8B model
Largest reduction versus prior DLM baselines
Smallest reduction versus prior DLM baselines
Math-specialized model result on AIME-24
Math-specialized model result on MATH-500
What the paper found
This paper from Texas A&M University introduces OPDLM, a data-efficient way to convert autoregressive language models into masked diffusion language models without full diffusion pretraining. The key idea is on-policy distillation: the student, initialized from a Qwen3 checkpoint and given bidirectional block-wise attention, generates its own reverse diffusion trajectories, and the frozen original ARLM provides token-level target logits on those same states. This directly addresses two bottlenecks in prior AR-to-diffusion conversion: knowledge loss from replacing next-token training with diffusion training, and the train-inference mismatch caused by training on random masks rather than sampler-induced states. On the main 8B setting, OPDLM trains with only 0.066B tokens and 4.2 × 10^18 FLOPs, yet remains competitive on benchmarks such as MMLU, GPQA-Diamond, AIME-24, and LiveCodeBench, while the full comparison set shows 15× to 7,000× fewer training tokens than established DLM baselines like SDAR, Dream, LLaDA, and Fast-dLLM-v2. The method also preserves zero-shot capabilities absent from its training corpus, including extended thinking and multilingual evaluation, and can be specialized for math: OPDLM-MATH-8B-Thinking reaches 93.8 on GSM8K, 92.4 on MATH-500, and 50.0 on AIME-24, all without a pretrained DLM stage or verifier reward. The work reframes ARLM-to-DLM conversion as post-training rather than pretraining.
Original abstract
We study the transformation of autoregressive models (ARLMs) into diffusion language models (DLMs). Rather than pretraining from scratch, prior work replaces the causal attention in ARLMs with bidirectional attention and then trains the resulting model using a DLM objective. However, these approaches incur two distribution shifts. First, transitioning from a next-token prediction objective to a DLM objective can discard knowledge acquired by the ARLM during training. Second, standard DLMs suffer from a train-inference mismatch, as the training loss is defined on randomly masked sequences rather than the trajectories encountered at inference produced by confidence-based decoding. To address both challenges, we introduce an On-Policy Diffusion Language Model (OPDLM) in which On-Policy Distillation (OPD) is employed for ARLM-to-DLM transformation. Specifically, OPDLM is trained via self-OPD, where the student, an ARLM with bidirectional attention, generates its own trajectories, and the teacher, the original frozen ARLM, distills its knowledge by providing target logits on these trajectories. By training directly in an on-policy manner, OPDLM eliminates the train-inference mismatch in DLMs, while distillation from the original model enhances knowledge retention from the ARLM. Empirical results demonstrate that OPDLM requires 15x to 7,000x fewer training tokens with strong performance across a wide variety of tasks. OPDLM avoids the prohibitive cost of DLM pretraining and positions DLM transformation as a form of ARLM post-training.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.