VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
AuthorsJiayi Yao, Samuel Shen, Kuntai Du, Shaoting Feng, Dongjoo Seo, Rui Zhang, Yuyang Huang, Yuhan Liu, Shan Lu, Junchen Jiang
Resources
VeriCache makes compressed LLM caching fast without changing the output, by drafting with a tiny cache and verifying against the full one only when needed.
Key results
On Qwen3-Coder-30B, KVzip 4× accumulates roughly this much sequence-level KL divergence per decoding step, showing why small per-step bias compounds over long outputs.
The paper reports that the per-step KL growth compounds to about this level after 250 decoding tokens for KVzip 4×, causing output drift from full-KV inference.
VeriCache’s compressed-KV drafting typically yields this many accepted tokens per verification round, far above typical small-model speculative drafters.
Used as the comparison point for VeriCache’s extended draft horizon; traditional speculative drafters usually accept only a few tokens per round.
Built on vLLM and LMCache, VeriCache achieves up to this throughput improvement on long-context decoding while preserving identical greedy-decoding outputs.
In the remote-prefix-caching setting, VeriCache reaches up to this speedup over full-KV inference while keeping outputs lossless.
What the paper found
VeriCache reframes KV-cache compression as a draft-and-verify accelerator rather than a replacement for exact inference: it runs the target LLM on a compressed KV cache to draft tokens, then verifies them against the full KV cache, so greedy-decoding output is identical to full-KV inference. The paper shows why this is necessary by measuring sequence-level KL divergence growth on Qwen3-Coder-30B, where KVzip 4× and TurboQuant k4v3 accumulate roughly 0.023 nats of KL per decoding step, which compounds to about 6 nats after 250 tokens and causes functional collapse on SWE-bench Lite and ComplexFuncBench even when token-level F1 remains above 75%. VeriCache’s key novelty is system design: cross-resource staggering overlaps compressed-KV drafting, which is HBM-bandwidth-bound, with full-KV reload and verification, which are interconnect- and compute-bound, and an extended draft horizon exploits unusually high acceptance lengths of 25–40 tokens per round, far above typical small-model speculative drafters’ 2–3. Built on vLLM and LMCache, with a uniform compressor interface spanning seven methods including KVzip, KIVI, KVQuant, and RotateKV, VeriCache reaches up to 4× higher throughput than full-KV inference on long-context decoding and up to 2× on remote prefix caching, while keeping KL below 0.01 nats from hardware nondeterminism and preserving exact quality on tasks like function calling, long-form generation, and code synthesis.
Original abstract
The large size of the KV cache has become a major bottleneck for serving LLMs with increasing context lengths. In response, many KV cache compression methods, such as token dropping and quantization, have been proposed. However, almost all of these methods are inherently lossy-despite minimal accuracy degradation for short outputs, their outputs increasingly diverge from full-KV-cache outputs as more tokens are decoded, which leads to catastrophic failures in code generation and tool calling. We present VeriCache, the first inference framework that ensures the same output as full-KV-cache decoding but largely preserves the high decoding throughput of a range of KV cache compression algorithms. VeriCache uses the compressed KV cache to draft tokens, then verifies them against the full KV cache. While it may seem like just speculative decoding, VeriCache requires addressing a key system challenge to work-keeping the full KV cache out of GPU memory and minimizing the overhead of swapping it in for verification. The insight is two-fold: (1) compressed-KV decoding can be parallelized with full-KV swap, because one is HBM-bandwidth-bound and the other is PCIe/network-bound, and (2) the compressed KV cache often produces output similar to the full KV cache, allowing a long drafting horizon to amortize each full-KV swap. VeriCache applies to both long-context decoding and remote prefix caching, supports a broad family of token-dropping and quantization methods through a uniform compressor interface, and composes with traditional speculative decoding. Experimental results show that VeriCache achieves up to 4X higher throughput than full-KV inference while producing identical outputs.
Read the original paper