Draft-OPD: On-Policy Distillation for Speculative Draft Models
AuthorsHaodi Lei, Yafu Li, Haoran Zhang, Shunkai Zhang, Qianjia Cheng, Xiaoye Qu, Ganqu Cui, Bowen Zhou, Ning Ding, Yun Luo, Yu Cheng
Resources
Draft-OPD improves speculative decoding by training draft models on the kinds of errors they actually make at inference time, yielding faster LLM generation with better acceptance rates.
Key results
Table 1 reports a mean speedup of 4.86× for Draft-OPD on Qwen3 models with thinking mode enabled and temperature 0.
Table 1 reports an average acceptance length of 6.60 tokens for Draft-OPD on Qwen3-4B in non-thinking mode at temperature 0.
What the paper found
Draft-OPD, from Shanghai Jiao Tong University and Shanghai AI Laboratory, tackles a bottleneck in speculative decoding: draft models such as EAGLE-3 and DFlash are usually trained with offline supervised fine-tuning on target-generated trajectories, but acceptance length plateaus because inference depends on draft-induced states, not static teacher data. The paper proposes on-policy distillation for draft models by combining target-assisted rollouts with error-position replay: the target model first produces a stable speculative rollout, anchors are recorded at each drafted block, and then the system replays drafting from those anchors so the target can score both accepted and rejected draft tokens at the exact failure states. Draft-OPD adds an acceptance-aware objective that uses forward KL on accepted tokens and reverse KL on rejected tokens, with exponentially decayed weights on later rejected positions to emphasize the earliest error in a block. Experiments on Qwen3-4B, Qwen3-8B, and Qwen3-30B-A3B-Thinking-2507 across GSM8K, MATH-500, AIME, MBPP, HumanEval, SWE-bench Lite, and MT-Bench show consistent gains over EAGLE-3 and DFlash under matched FLOPs: average speedup rises to 4.86× on thinking-mode models and 5.31× without thinking, with average acceptance lengths up to 6.60 tokens in non-thinking settings. On SGLang with FA3, Draft-OPD improves throughput by 6% to 17% depending on model and task, confirming that verification-time error replay translates into real serving gains.
Original abstract
Speculative decoding accelerates large language model inference by pairing a target model with a lightweight draft model whose proposed tokens are verified in parallel. A common way to build draft models, like EAGLE3 or DFlash is supervised fine-tuning (SFT) on target-generated trajectories. However, we observe that SFT quickly plateaus: the draft model's acceptance length on test data stops improving. The reason is an offline-to-inference mismatch: In SFT, the drafter learns from fixed target-generated trajectories, whereas during speculative decoding it is evaluated on blocks proposed under its own policy. This motivates on-policy distillation (OPD), where the target model supervises the drafter on draft-induced states. Yet OPD remains difficult for draft models, as they cannot reliably roll out complete sequences independently, whereas target-assisted generation makes the collected sequences follow the target distribution and thus eliminates the on-policy signal. We therefore propose Draft-OPD, which uses target-assisted rollout for stable continuations and replays drafting from the verification-exposed error positions. This allows the drafter to learn from target feedback on both accepted and rejected proposals, focusing training on the draft-induced errors that limit speculative acceptance. Experiments show that Draft-OPD achieves over $5\times$ lossless acceleration for thinking models across diverse tasks, improving over EAGLE-3 and DFlash by 23\% and 13\%.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.