NTH
AI research

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference

AuthorsJiayi Yao, Samuel Shen, Kuntai Du, Shaoting Feng, Dongjoo Seo, Rui Zhang, Yuyang Huang, Yuhan Liu, Shan Lu, Junchen Jiang

May 25, 2026 3 min read
Watch on YouTube
The one-line take

VeriCache makes compressed LLM caching fast without changing the output, by drafting with a tiny cache and verifying against the full one only when needed.

Key results

0.023 nats
KL growth per decoding step

On Qwen3-Coder-30B, KVzip 4× accumulates roughly this much sequence-level KL divergence per decoding step, showing why small per-step bias compounds over long outputs.

6 nats
KL after 250 tokens

The paper reports that the per-step KL growth compounds to about this level after 250 decoding tokens for KVzip 4×, causing output drift from full-KV inference.

25–40 tokens
Acceptance length per verification round

VeriCache’s compressed-KV drafting typically yields this many accepted tokens per verification round, far above typical small-model speculative drafters.

2–3 tokens
Typical small-model drafter acceptance

Used as the comparison point for VeriCache’s extended draft horizon; traditional speculative drafters usually accept only a few tokens per round.

up to 4×
Throughput gain over full-KV inference

Built on vLLM and LMCache, VeriCache achieves up to this throughput improvement on long-context decoding while preserving identical greedy-decoding outputs.

up to 2×
Remote prefix caching speedup

In the remote-prefix-caching setting, VeriCache reaches up to this speedup over full-KV inference while keeping outputs lossless.

What the paper found

VeriCache reframes KV-cache compression as a draft-and-verify accelerator rather than a replacement for exact inference: it runs the target LLM on a compressed KV cache to draft tokens, then verifies them against the full KV cache, so greedy-decoding output is identical to full-KV inference. The paper shows why this is necessary by measuring sequence-level KL divergence growth on Qwen3-Coder-30B, where KVzip 4× and TurboQuant k4v3 accumulate roughly 0.023 nats of KL per decoding step, which compounds to about 6 nats after 250 tokens and causes functional collapse on SWE-bench Lite and ComplexFuncBench even when token-level F1 remains above 75%. VeriCache’s key novelty is system design: cross-resource staggering overlaps compressed-KV drafting, which is HBM-bandwidth-bound, with full-KV reload and verification, which are interconnect- and compute-bound, and an extended draft horizon exploits unusually high acceptance lengths of 25–40 tokens per round, far above typical small-model speculative drafters’ 2–3. Built on vLLM and LMCache, with a uniform compressor interface spanning seven methods including KVzip, KIVI, KVQuant, and RotateKV, VeriCache reaches up to 4× higher throughput than full-KV inference on long-context decoding and up to 2× on remote prefix caching, while keeping KL below 0.01 nats from hardware nondeterminism and preserving exact quality on tasks like function calling, long-form generation, and code synthesis.

Original abstract

The large size of the KV cache has become a major bottleneck for serving LLMs with increasing context lengths. In response, many KV cache compression methods, such as token dropping and quantization, have been proposed. However, almost all of these methods are inherently lossy-despite minimal accuracy degradation for short outputs, their outputs increasingly diverge from full-KV-cache outputs as more tokens are decoded, which leads to catastrophic failures in code generation and tool calling. We present VeriCache, the first inference framework that ensures the same output as full-KV-cache decoding but largely preserves the high decoding throughput of a range of KV cache compression algorithms. VeriCache uses the compressed KV cache to draft tokens, then verifies them against the full KV cache. While it may seem like just speculative decoding, VeriCache requires addressing a key system challenge to work-keeping the full KV cache out of GPU memory and minimizing the overhead of swapping it in for verification. The insight is two-fold: (1) compressed-KV decoding can be parallelized with full-KV swap, because one is HBM-bandwidth-bound and the other is PCIe/network-bound, and (2) the compressed KV cache often produces output similar to the full KV cache, allowing a long drafting horizon to amortize each full-KV swap. VeriCache applies to both long-context decoding and remote prefix caching, supports a broad family of token-dropping and quantization methods through a uniform compressor interface, and composes with traditional speculative decoding. Experimental results show that VeriCache achieves up to 4X higher throughput than full-KV inference while producing identical outputs.

Read the original paper