NTH

Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy Distillation

AuthorsXingyu Su, Jacob Helwig, Shubham Parashar, Atharv Chagi, Lakshmi Jotsna, Degui Zhi, James Caverlee, Dileep Kalathil, Shuiwang Ji

July 9, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows how to turn an autoregressive language model into a diffusion-style language model much more efficiently by training it on its own generated trajectories while distilling knowledge from the original model.

Key results

0.066B
OPDLM-8B training tokens

Training data used for the main 8B model

4.2
OPDLM-8B FLOPs

Training compute in 10^18 FLOPs for the main 8B model

7000x
Training token reduction

Largest reduction versus prior DLM baselines

15x
Training token reduction

Smallest reduction versus prior DLM baselines

50.0
OPDLM-MATH-8B-Thinking AIME-24

Math-specialized model result on AIME-24

92.4
OPDLM-MATH-8B-Thinking MATH-500

Math-specialized model result on MATH-500

What the paper found

This paper from Texas A&M University introduces OPDLM, a data-efficient way to convert autoregressive language models into masked diffusion language models without full diffusion pretraining. The key idea is on-policy distillation: the student, initialized from a Qwen3 checkpoint and given bidirectional block-wise attention, generates its own reverse diffusion trajectories, and the frozen original ARLM provides token-level target logits on those same states. This directly addresses two bottlenecks in prior AR-to-diffusion conversion: knowledge loss from replacing next-token training with diffusion training, and the train-inference mismatch caused by training on random masks rather than sampler-induced states. On the main 8B setting, OPDLM trains with only 0.066B tokens and 4.2 × 10^18 FLOPs, yet remains competitive on benchmarks such as MMLU, GPQA-Diamond, AIME-24, and LiveCodeBench, while the full comparison set shows 15× to 7,000× fewer training tokens than established DLM baselines like SDAR, Dream, LLaDA, and Fast-dLLM-v2. The method also preserves zero-shot capabilities absent from its training corpus, including extended thinking and multilingual evaluation, and can be specialized for math: OPDLM-MATH-8B-Thinking reaches 93.8 on GSM8K, 92.4 on MATH-500, and 50.0 on AIME-24, all without a pretrained DLM stage or verifier reward. The work reframes ARLM-to-DLM conversion as post-training rather than pretraining.

Original abstract

We study the transformation of autoregressive models (ARLMs) into diffusion language models (DLMs). Rather than pretraining from scratch, prior work replaces the causal attention in ARLMs with bidirectional attention and then trains the resulting model using a DLM objective. However, these approaches incur two distribution shifts. First, transitioning from a next-token prediction objective to a DLM objective can discard knowledge acquired by the ARLM during training. Second, standard DLMs suffer from a train-inference mismatch, as the training loss is defined on randomly masked sequences rather than the trajectories encountered at inference produced by confidence-based decoding. To address both challenges, we introduce an On-Policy Diffusion Language Model (OPDLM) in which On-Policy Distillation (OPD) is employed for ARLM-to-DLM transformation. Specifically, OPDLM is trained via self-OPD, where the student, an ARLM with bidirectional attention, generates its own trajectories, and the teacher, the original frozen ARLM, distills its knowledge by providing target logits on these trajectories. By training directly in an on-policy manner, OPDLM eliminates the train-inference mismatch in DLMs, while distillation from the original model enhances knowledge retention from the ARLM. Empirical results demonstrate that OPDLM requires 15x to 7,000x fewer training tokens with strong performance across a wide variety of tasks. OPDLM avoids the prohibitive cost of DLM pretraining and positions DLM transformation as a form of ARLM post-training.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis