NTH

JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting

AuthorsLanxiang Hu, Zhaoxiang Feng, Yulun Wu, Haoran Yuan, Yujie Zhao, Yu-Yang Qian, Bojun Wang, Peng Zhao, Daxin Jiang, Yibo Zhu, Tajana Rosing, Hao Zhang

June 29, 2026 2 min read
Watch on YouTube
The one-line take

JetSpec makes large language models decode faster by drafting and verifying better token trees in parallel, pushing speculative decoding to much higher speeds.

Key results

780K
training mixture

examples from Nemotron Post-Training Dataset V2

9.64×
speedup MATH-500

JetSpec on Qwen3-8B with 256-token tree budget

4.58×
speedup MT-Bench

JetSpec on Qwen3-8B with 256-token tree budget

9.45×
speedup Qwen3-30B-A3B MATH-500

JetSpec on the MoE target model with 256-token tree budget

553.3
throughput

vLLM serving on a single H100 at batch size 1 with JetSpec budget 128, in TPS

What the paper found

JetSpec, from UC San Diego, Zhejiang University, UIUC, Nanjing University, and StepFun, attacks the speculative decoding scaling ceiling by combining one-pass parallel drafting with branch-wise causal conditioning. Instead of DFlash-style block diffusion, which can generate mutually inconsistent candidate trees, JetSpec trains a causal parallel draft head on fused hidden states from a frozen target model so each node is conditioned on its own ancestor path. The paper formalizes the speedup bottleneck in terms of draft length, acceptance rate, and per-token draft cost, then shows that a tree-causal attention mask aligns draft scores with the target model’s autoregressive factorization. Trained on 780K examples from NVIDIA’s Nemotron Post-Training Dataset V2 plus 20K CodeAlpaca examples, JetSpec is evaluated on Qwen3-8B and Qwen3-30B-A3B across GSM8K, MATH-500, AIME25, HumanEval, MBPP, LiveCodeBench, and MT-Bench. On H100 GPUs, it reaches 9.64× speedup on MATH-500 and 4.58× on MT-Bench with a 256-token tree budget, outperforming EAGLE-3, DFlash, and DDTree; on Qwen3-30B-A3B it still achieves 9.45× on MATH-500. In vLLM serving, a single H100 sees throughput rise from 127.8 TPS to 553.3 TPS at batch size 1 with budget 128, showing the method’s practical latency gains under realistic load.

Original abstract

Speculative decoding (SD) accelerates autoregressive Large Language Models (LLMs) by drafting multiple tokens and verifying them in parallel, but it faces a scaling limitation: increasing the draft budget improves speed only when acceptance remains high and drafting overhead stays low. This ceiling has been difficult to break because prior head-based SD methods face a causality-efficiency dilemma. Autoregressive drafters produce path-conditioned candidates that are effective for tree speculative decoding with higher acceptance length, but their drafting cost grows with tree depth. Bidirectional block-diffusion drafters generate all positions in one pass, but their branch-agnostic marginals can form individually plausible yet mutually inconsistent trees, wasting budget and reducing acceptance. We propose JetSpec, a head-based SD framework that combines one-forward drafting efficiency with branch-wise causal conditioning. JetSpec trains a causal parallel draft head over fused hidden states from the frozen target model, producing candidate trees whose scores align with the target model's autoregressive factorization. This enables JetSpec to convert larger draft budgets into longer accepted prefixes and higher end-to-end speedup. Across math, coding, and chat benchmarks on dense and MoE Qwen3 models, JetSpec consistently outperforms bidirectional-head and tree-based SD baselines. On H100 GPUs, JetSpec achieves up to 9.64x speedup on MATH-500 and 4.58x on open-ended conversational workloads, with further latency gains demonstrated through vLLM integration under realistic serving loads. Our code and models are available at https://github.com/hao-ai-lab/JetSpec.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis