NTH

Draft-OPD: On-Policy Distillation for Speculative Draft Models

AuthorsHaodi Lei, Yafu Li, Haoran Zhang, Shunkai Zhang, Qianjia Cheng, Xiaoye Qu, Ganqu Cui, Bowen Zhou, Ning Ding, Yun Luo, Yu Cheng

June 7, 2026 2 min read
Watch on YouTube
The one-line take

Draft-OPD improves speculative decoding by training draft models on the kinds of errors they actually make at inference time, yielding faster LLM generation with better acceptance rates.

Key results

4.86×
Mean speedup, thinking mode

Table 1 reports a mean speedup of 4.86× for Draft-OPD on Qwen3 models with thinking mode enabled and temperature 0.

6.60
Mean acceptance length, non-thinking mode

Table 1 reports an average acceptance length of 6.60 tokens for Draft-OPD on Qwen3-4B in non-thinking mode at temperature 0.

What the paper found

Draft-OPD, from Shanghai Jiao Tong University and Shanghai AI Laboratory, tackles a bottleneck in speculative decoding: draft models such as EAGLE-3 and DFlash are usually trained with offline supervised fine-tuning on target-generated trajectories, but acceptance length plateaus because inference depends on draft-induced states, not static teacher data. The paper proposes on-policy distillation for draft models by combining target-assisted rollouts with error-position replay: the target model first produces a stable speculative rollout, anchors are recorded at each drafted block, and then the system replays drafting from those anchors so the target can score both accepted and rejected draft tokens at the exact failure states. Draft-OPD adds an acceptance-aware objective that uses forward KL on accepted tokens and reverse KL on rejected tokens, with exponentially decayed weights on later rejected positions to emphasize the earliest error in a block. Experiments on Qwen3-4B, Qwen3-8B, and Qwen3-30B-A3B-Thinking-2507 across GSM8K, MATH-500, AIME, MBPP, HumanEval, SWE-bench Lite, and MT-Bench show consistent gains over EAGLE-3 and DFlash under matched FLOPs: average speedup rises to 4.86× on thinking-mode models and 5.31× without thinking, with average acceptance lengths up to 6.60 tokens in non-thinking settings. On SGLang with FA3, Draft-OPD improves throughput by 6% to 17% depending on model and task, confirming that verification-time error replay translates into real serving gains.

Original abstract

Speculative decoding accelerates large language model inference by pairing a target model with a lightweight draft model whose proposed tokens are verified in parallel. A common way to build draft models, like EAGLE3 or DFlash is supervised fine-tuning (SFT) on target-generated trajectories. However, we observe that SFT quickly plateaus: the draft model's acceptance length on test data stops improving. The reason is an offline-to-inference mismatch: In SFT, the drafter learns from fixed target-generated trajectories, whereas during speculative decoding it is evaluated on blocks proposed under its own policy. This motivates on-policy distillation (OPD), where the target model supervises the drafter on draft-induced states. Yet OPD remains difficult for draft models, as they cannot reliably roll out complete sequences independently, whereas target-assisted generation makes the collected sequences follow the target distribution and thus eliminates the on-policy signal. We therefore propose Draft-OPD, which uses target-assisted rollout for stable continuations and replays drafting from the verification-exposed error positions. This allows the drafter to learn from target feedback on both accepted and rejected proposals, focusing training on the draft-induced errors that limit speculative acceptance. Experiments show that Draft-OPD achieves over $5\times$ lossless acceleration for thinking models across diverse tasks, improving over EAGLE-3 and DFlash by 23\% and 13\%.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis