JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting
AuthorsLanxiang Hu, Zhaoxiang Feng, Yulun Wu, Haoran Yuan, Yujie Zhao, Yu-Yang Qian, Bojun Wang, Peng Zhao, Daxin Jiang, Yibo Zhu, Tajana Rosing, Hao Zhang
JetSpec makes large language models decode faster by drafting and verifying better token trees in parallel, pushing speculative decoding to much higher speeds.
Key results
examples from Nemotron Post-Training Dataset V2
JetSpec on Qwen3-8B with 256-token tree budget
JetSpec on Qwen3-8B with 256-token tree budget
JetSpec on the MoE target model with 256-token tree budget
vLLM serving on a single H100 at batch size 1 with JetSpec budget 128, in TPS
What the paper found
JetSpec, from UC San Diego, Zhejiang University, UIUC, Nanjing University, and StepFun, attacks the speculative decoding scaling ceiling by combining one-pass parallel drafting with branch-wise causal conditioning. Instead of DFlash-style block diffusion, which can generate mutually inconsistent candidate trees, JetSpec trains a causal parallel draft head on fused hidden states from a frozen target model so each node is conditioned on its own ancestor path. The paper formalizes the speedup bottleneck in terms of draft length, acceptance rate, and per-token draft cost, then shows that a tree-causal attention mask aligns draft scores with the target model’s autoregressive factorization. Trained on 780K examples from NVIDIA’s Nemotron Post-Training Dataset V2 plus 20K CodeAlpaca examples, JetSpec is evaluated on Qwen3-8B and Qwen3-30B-A3B across GSM8K, MATH-500, AIME25, HumanEval, MBPP, LiveCodeBench, and MT-Bench. On H100 GPUs, it reaches 9.64× speedup on MATH-500 and 4.58× on MT-Bench with a 256-token tree budget, outperforming EAGLE-3, DFlash, and DDTree; on Qwen3-30B-A3B it still achieves 9.45× on MATH-500. In vLLM serving, a single H100 sees throughput rise from 127.8 TPS to 553.3 TPS at batch size 1 with budget 128, showing the method’s practical latency gains under realistic load.
Original abstract
Speculative decoding (SD) accelerates autoregressive Large Language Models (LLMs) by drafting multiple tokens and verifying them in parallel, but it faces a scaling limitation: increasing the draft budget improves speed only when acceptance remains high and drafting overhead stays low. This ceiling has been difficult to break because prior head-based SD methods face a causality-efficiency dilemma. Autoregressive drafters produce path-conditioned candidates that are effective for tree speculative decoding with higher acceptance length, but their drafting cost grows with tree depth. Bidirectional block-diffusion drafters generate all positions in one pass, but their branch-agnostic marginals can form individually plausible yet mutually inconsistent trees, wasting budget and reducing acceptance. We propose JetSpec, a head-based SD framework that combines one-forward drafting efficiency with branch-wise causal conditioning. JetSpec trains a causal parallel draft head over fused hidden states from the frozen target model, producing candidate trees whose scores align with the target model's autoregressive factorization. This enables JetSpec to convert larger draft budgets into longer accepted prefixes and higher end-to-end speedup. Across math, coding, and chat benchmarks on dense and MoE Qwen3 models, JetSpec consistently outperforms bidirectional-head and tree-based SD baselines. On H100 GPUs, JetSpec achieves up to 9.64x speedup on MATH-500 and 4.58x on open-ended conversational workloads, with further latency gains demonstrated through vLLM integration under realistic serving loads. Our code and models are available at https://github.com/hao-ai-lab/JetSpec.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.