NTH

ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression

AuthorsYuhang Zhan, Lisi Chen, Shuo Shang

August 8, 2026 2 min read
Watch on YouTube
The one-line take

ResKV makes long-context language-model decoding more efficient by compressing KV caches while reconstructing the attention contributions of discarded tokens.

Key results

32
LongBench configurations improved

ResKV improves all 32 displayed LongBench configurations.

1.02
LongBench average gain

Average score improvement across LLaMA-3.1-8B-Instruct and Qwen-2.5-7B-Instruct.

63
RULER configurations improved

ResKV improves 63 of 64 displayed RULER configurations.

3.38
RULER average gain

Average improvement across 4K and 32K contexts and both backbones.

4.54
Query-agnostic RULER gain

Average improvement for query-agnostic construction, versus 2.22 for query-aware construction.

3.47
Tight-budget LongBench gain

Improvement at 10% retained KV; the gain is 3.66 at 20% retained KV.

What the paper found

ResKV, from Yuhang Zhan, Lisi Chen, and Shuo Shang at the University of Electronic Science and Technology of China, addresses a weakness in KV-cache compression: hard eviction discards both the numerator and denominator contributions of omitted tokens, while merging methods can corrupt retained memories. Its fixed-budget design splits the cache into an exact main cache and compact residual entries, each representing a cluster with a mean key, mean value, and population count. Through a shared softmax, residual entries restore omitted attention mass inside the same normalization rather than applying a post-hoc output correction. A construction-time validation proxy allocates residual slots per layer and KV head, while a decode-time dynamic gate suppresses residual influence when the main-cache attention is sharply concentrated. Experiments on LLaMA-3.1-8B-Instruct and Qwen-2.5-7B-Instruct, using LongBench and RULER, show improvements under identical retained KV budgets: ResKV improves all 32 displayed LongBench configurations with an average gain of 1.02 points, and improves 63 of 64 displayed RULER configurations with an average gain of 3.38 points. In query-agnostic construction, the average RULER improvement reaches 4.54 points, compared with 2.22 points for query-aware construction, while tight LongBench budgets produce gains of 3.47 points at 10% retained KV and 3.66 points at 20% retained KV. The method preserves the compressed baseline’s peak memory footprint and stable long-context decoding, although its residual branch adds moderate throughput cost.

Original abstract

KV cache compression is essential for efficient long-context inference. Existing eviction methods permanently discard unselected tokens and consequently remove their aggregate contribution to attention. Merging-based alternatives preserve more information but can perturb retained keys and values that should remain exact. We observe that the information omitted by cache eviction can be formulated as residual statistics in both the numerator and denominator of softmax attention. Based on this observation, we propose ResKV, which divides a fixed KV budget into an exact main cache and a compact residual cache that reconstructs the contribution of omitted tokens. ResKV lets main-cache tokens and residual entries participate in the same softmax normalization, so residual entries restore both attention numerator and denominator mass rather than acting as a post-hoc correction. A construction-time validation proxy determines residual allocation for each layer and KV head, while a decode-time dynamic gate adjusts residual contributions for individual queries. Comprehensive evaluations on LongBench and RULER, covering query-aware and query-agnostic settings, multiple backbones, cache budgets, and representative compression baselines, demonstrate broad improvements under the same retained KV budget while preserving the practical efficiency of compressed decoding, including peak memory usage and long-context decode throughput.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis