NTH

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

AuthorsXin Cheng, Xingkai Yu, Chenze Shao, Jiashi Li, Yunfan Xiong, Yi Qian, Jiaqi Zhu, Shirong Ma, Xiaokang Zhang, Jiasheng Ye, Qinyu Chen, Chengqi Deng, Jiping Yu, Damai Dai, Zhengyan Zhang, Yixuan Wei, Yixuan Tan, Wenkai Yang, Runxin Xu, Yu Wu, Zhean Xu, Xuanyu Wang, Muyang Chen, Rui Tian, Xiao Bi, Zhewen Hao, Shaoyuan Chen, Huanqi Cao, Wentao Zhang, Anyi Xu, Huishuai Zhang, Dongyan Zhao, Wenfeng Liang

July 11, 2026 2 min read
Watch on YouTube
The one-line take

DSpark speeds up LLM inference by generating draft tokens in a more structured way and verifying them adaptively, delivering major serving throughput gains in real production traffic.

Key results

30.9%
Qwen3-4B gain vs Eagle3

Macro-average accepted length improvement on offline benchmarks

26.7%
Qwen3-8B gain vs Eagle3

Macro-average accepted length improvement on offline benchmarks

30.0%
Qwen3-14B gain vs Eagle3

Macro-average accepted length improvement on offline benchmarks

60%
DeepSeek-V4-Flash speedup

Per-user generation speed improvement at matched throughput in production

57%
DeepSeek-V4-Pro speedup

Per-user generation speed improvement at matched throughput in production

What the paper found

DSpark, from DeepSeek-AI and Peking University, is a speculative decoding framework that combines a semi-autoregressive drafter with confidence-scheduled verification to solve two production bottlenecks at once: suffix decay in parallel draft generation and wasted verification capacity under high concurrency. Instead of relying on a fully independent parallel block like DFlash or a fully sequential drafter like Eagle3, DSpark keeps a heavy parallel backbone and adds a lightweight sequential Markov or RNN head to inject local token dependence while preserving single-pass draft latency. It then trains a confidence head to estimate prefix survival probabilities and calibrates those scores with Sequential Temperature Scaling so they can drive a hardware-aware prefix scheduler that maximizes expected throughput using real engine load profiles. On offline benchmarks across Qwen3-4B, 8B, 14B, and Gemma4-12B, evaluated on GSM8K, MATH500, AIME25, MBPP, HumanEval, Live-CodeBench, MT-Bench, Alpaca, and Arena-Hard, DSpark raises accepted length over Eagle3 by 30.9%, 26.7%, and 30.0% on the three Qwen3 scales, and over DFlash by 16.3%, 18.4%, and 18.3%. In DeepSeek-V4 production traffic, compared with the MTP-1 baseline, DSpark improves per-user generation speed by 60%–85% on V4-Flash and 57%–78% on V4-Pro at matched throughput, while under strict SLAs it preserves usable capacity where the baseline collapses, shifting the serving Pareto frontier.

Original abstract

Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification. While recent parallel drafters efficiently propose long token sequences in a single forward pass, they suffer from rapid acceptance decay due to a lack of inter-token dependencies. Furthermore, indiscriminately verifying these extended blocks wastes critical batch capacity on tokens with high rejection risks, severely degrading throughput in high-concurrency serving systems. We introduce DSpark, a speculative decoding framework that unifies high-throughput parallel generation with adaptive, load-aware verification. To maintain draft quality, DSpark utilizes a semi-autoregressive architecture, coupling a parallel backbone with a lightweight sequential module, to introduce intra-block dependency modeling and mitigate suffix decay. To optimize system efficiency, DSpark employs confidence-scheduled verification, dynamically tailoring the verification length for each request based on estimated prefix survival probabilities and engine-specific throughput profiles. On offline benchmarks across diverse domains, DSpark substantially improves the accepted length over state-of-the-art autoregressive and parallel drafters. When deployed within the DeepSeek-V4 serving system under live user traffic, DSpark successfully mitigates verification waste. Compared to the established production baseline (MTP-1), DSpark accelerates per-user generation speeds by 60 to 85 percent at matched throughput levels. More importantly, by preventing severe throughput degradation under strict interactivity constraints, it enables performance tiers that were previously unattainable, shifting the Pareto frontier of our serving system.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis