NTH

Disaggregated Quantization: Specializing LLM Prefill and Decode

AuthorsAndrei Panferov, Maximilian Kleinegger, Sweta Priyadarshi, Tijmen Blankevoort, Dan Alistarh

AffiliationsNVIDIA · ISTA

October 1, 2026 2 min read
Watch on YouTube
The one-line take

Disaggregated quantization gives LLM prefill and decode their own specialized weights and formats, improving low-bit accuracy while speeding up first-token generation.

Key results

32.5
MMLU-Pro gain

Percentage-point accuracy improvement from an NVFP4 prefiller for the 1-bit Qwen3.8-27B decoder.

35.3
MMMU-Pro gain

Percentage-point accuracy improvement from the same NVFP4 prefiller on visual reasoning.

1.78
ODP TTFT speedup

Time-to-first-token speedup over weight-only inference at 8K context in llama.cpp.

2.8T
Largest PTQ model

Maximum model scale used for post-training validation of format disaggregation.

What the paper found

This paper introduces disaggregated quantization, or DQ, which treats LLM prefill and decode as different computational workloads: prefill is compute-bound and benefits from NVFP4 weight-and-activation quantization, while decode is memory-bound and benefits from compact weight-only formats such as LUT2 and LUT3. Its quantization-aware distillation with disaggregation, QADD, uses the instruction-tuning mask to route prompt tokens through a prefill pathway and response tokens through a decode pathway while optimizing both against a frozen BF16 teacher. Experiments on Qwen 3 and Gemma 3 show that disabling activation quantization during decode improves decode-heavy accuracy without increasing storage or inference cost, while fully separate prefill weights improve low-bit accuracy on both workload types. For an existing Qwen3.8-27B GGUF decoder, an NVFP4 prefiller raises 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 points on MMMU-Pro without changing the decode checkpoint. Offloaded disaggregated prefill, or ODP, streams the extra prefiller from SSD and overlaps loading with computation; in llama.cpp, it delivers a 1.78× time-to-first-token speedup at 8K context. The method was also validated with post-training quantization on models up to 2.8T parameters, including Qwen, Gemma, Meta’s Muse Glimmer, NVIDIA Nemotron, and Kimi models.

Original abstract

Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and storage placement to both of these phases. On Qwen 3 and Gemma 3, removing activation quantization specifically on decode improves accuracy on decode-heavy tasks without increasing inference cost. Training separate compute-native prefill weights accelerates prompt processing relative to weight-only inference while matching or exceeding its accuracy at 2-3-bit decode on both decode-heavy and prefill-heavy tasks. With released Qwen3.8-27B GGUF decoders, training an NVFP4 prefiller improves 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without modifying the decode checkpoint. To accommodate the additional checkpoint on a single device, offloaded disaggregated prefill (ODP) streams its weights from SSD, amortizing loading over prompt length. On the same 27B model, ODP delivers a 1.78x time-to-first-token speedup over the weight-only baseline at 8K prompt length in llama.cpp. We evaluate accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis