NTH

Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding

AuthorsJianuo Huang, Yaojie Zhang, Qituan Zhang, Hao Lin, Hanlin Xu, Linfeng Zhang

June 7, 2026 2 min read
Watch on YouTube
The one-line take

Domino speeds up LLM generation by using a parallel draft model plus a lightweight causal refinement step, delivering faster speculative decoding without sacrificing too much draft quality.

Key results

5.49
Qwen3-8B avg speedup

Domino reaches 5.49× average end-to-end speedup on Qwen3-8B under the Transformers backend.

5.8
SGLang throughput speedup

Domino achieves up to 5.8× throughput speedup under SGLang serving.

7.92
GSM8K speedup

On GSM8K, Domino achieves 7.92× speedup on Qwen3-8B, compared with 5.21× for DFlash.

4.19
Domino head accept length

Ablation shows the Domino head raises average acceptance length from 3.49 to 4.19.

3.31
Domino head speedup

Ablation shows the Domino head raises average speedup from 2.84× to 3.31×.

What the paper found

Domino, from EPIC Lab at Shanghai Jiao Tong University and Huawei collaborators, tackles a core bottleneck in speculative decoding for large language models such as Qwen3-4B and Qwen3-8B: autoregressive drafters like EAGLE-3 preserve causal token dependencies but pay a sequential cost, while fully parallel drafters like DFlash cut latency but weaken intra-block dependency modeling. Domino decouples these two goals by using a parallel draft backbone to generate block-level preliminary distributions in one pass, then adding a lightweight causal correction branch with a GRU-based encoder and a low-rank logit-space head to inject prefix information without re-running a full LM head. The training recipe is equally novel: a teacher-forced causal encoder avoids noisy self-generated prefixes, and a base-anchored curriculum first optimizes the parallel backbone before gradually shifting weight to the corrected final logits. On math, code, and chat benchmarks including GSM8K, MATH-500, AIME25, HumanEval, MBPP, LiveCodeBench, MT-Bench, and Alpaca, Domino improves average acceptance length and speedup over representative baselines; on Qwen3-8B it reaches 5.49× average end-to-end speedup under the Transformers backend and up to 5.8× throughput speedup in SGLang, with a standout 7.92× speedup on GSM8K versus 5.21× for DFlash. Ablations show the Domino head alone raises average acceptance length from 3.49 to 4.19 and speedup from 2.84× to 3.31×, confirming that lightweight causal correction is the key mechanism.

Original abstract

Speculative decoding accelerates LLM inference by drafting multiple tokens and verifying them in parallel with the target model. However, its practical speedup is constrained by the trade-off between draft quality and drafting cost: autoregressive drafters model causal dependencies among draft tokens but incur sequential overhead, while parallel drafters reduce drafting cost but weaken intra-block dependency modeling. In this paper, we propose Domino, a speculative decoding framework that decouples causal dependency modeling from expensive autoregressive draft execution. Domino first uses a parallel draft backbone to produce preliminary draft distributions for the entire block, and then applies a lightweight Domino head to refine them with prefix-dependent causal information. To stabilize teacher-forced causal encoding, we further introduce a base-anchored training curriculum that first strengthens the parallel backbone and then gradually shifts optimization toward the causally corrected final distribution. Experiments on Qwen3 models show that Domino achieves up to \(5.49\times\) end-to-end speedup under the Transformers backend and up to \(5.8\times\) throughput speedup under SGLang serving.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis