Disaggregated Quantization: Specializing LLM Prefill and Decode
AuthorsAndrei Panferov, Maximilian Kleinegger, Sweta Priyadarshi, Tijmen Blankevoort, Dan Alistarh
AffiliationsNVIDIA · ISTA
Resources
Disaggregated quantization gives LLM prefill and decode their own specialized weights and formats, improving low-bit accuracy while speeding up first-token generation.
Key results
Percentage-point accuracy improvement from an NVFP4 prefiller for the 1-bit Qwen3.8-27B decoder.
Percentage-point accuracy improvement from the same NVFP4 prefiller on visual reasoning.
Time-to-first-token speedup over weight-only inference at 8K context in llama.cpp.
Maximum model scale used for post-training validation of format disaggregation.
What the paper found
This paper introduces disaggregated quantization, or DQ, which treats LLM prefill and decode as different computational workloads: prefill is compute-bound and benefits from NVFP4 weight-and-activation quantization, while decode is memory-bound and benefits from compact weight-only formats such as LUT2 and LUT3. Its quantization-aware distillation with disaggregation, QADD, uses the instruction-tuning mask to route prompt tokens through a prefill pathway and response tokens through a decode pathway while optimizing both against a frozen BF16 teacher. Experiments on Qwen 3 and Gemma 3 show that disabling activation quantization during decode improves decode-heavy accuracy without increasing storage or inference cost, while fully separate prefill weights improve low-bit accuracy on both workload types. For an existing Qwen3.8-27B GGUF decoder, an NVFP4 prefiller raises 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 points on MMMU-Pro without changing the decode checkpoint. Offloaded disaggregated prefill, or ODP, streams the extra prefiller from SSD and overlaps loading with computation; in llama.cpp, it delivers a 1.78× time-to-first-token speedup at 8K context. The method was also validated with post-training quantization on models up to 2.8T parameters, including Qwen, Gemma, Meta’s Muse Glimmer, NVIDIA Nemotron, and Kimi models.
Original abstract
Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and storage placement to both of these phases. On Qwen 3 and Gemma 3, removing activation quantization specifically on decode improves accuracy on decode-heavy tasks without increasing inference cost. Training separate compute-native prefill weights accelerates prompt processing relative to weight-only inference while matching or exceeding its accuracy at 2-3-bit decode on both decode-heavy and prefill-heavy tasks. With released Qwen3.8-27B GGUF decoders, training an NVFP4 prefiller improves 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without modifying the decode checkpoint. To accommodate the additional checkpoint on a single device, offloaded disaggregated prefill (ODP) streams its weights from SSD, amortizing loading over prompt length. On the same 27B model, ODP delivers a 1.78x time-to-first-token speedup over the weight-only baseline at 8K prompt length in llama.cpp. We evaluate accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.