ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression
AuthorsYuhang Zhan, Lisi Chen, Shuo Shang
Resources
ResKV makes long-context language-model decoding more efficient by compressing KV caches while reconstructing the attention contributions of discarded tokens.
Key results
ResKV improves all 32 displayed LongBench configurations.
Average score improvement across LLaMA-3.1-8B-Instruct and Qwen-2.5-7B-Instruct.
ResKV improves 63 of 64 displayed RULER configurations.
Average improvement across 4K and 32K contexts and both backbones.
Average improvement for query-agnostic construction, versus 2.22 for query-aware construction.
Improvement at 10% retained KV; the gain is 3.66 at 20% retained KV.
What the paper found
ResKV, from Yuhang Zhan, Lisi Chen, and Shuo Shang at the University of Electronic Science and Technology of China, addresses a weakness in KV-cache compression: hard eviction discards both the numerator and denominator contributions of omitted tokens, while merging methods can corrupt retained memories. Its fixed-budget design splits the cache into an exact main cache and compact residual entries, each representing a cluster with a mean key, mean value, and population count. Through a shared softmax, residual entries restore omitted attention mass inside the same normalization rather than applying a post-hoc output correction. A construction-time validation proxy allocates residual slots per layer and KV head, while a decode-time dynamic gate suppresses residual influence when the main-cache attention is sharply concentrated. Experiments on LLaMA-3.1-8B-Instruct and Qwen-2.5-7B-Instruct, using LongBench and RULER, show improvements under identical retained KV budgets: ResKV improves all 32 displayed LongBench configurations with an average gain of 1.02 points, and improves 63 of 64 displayed RULER configurations with an average gain of 3.38 points. In query-agnostic construction, the average RULER improvement reaches 4.54 points, compared with 2.22 points for query-aware construction, while tight LongBench budgets produce gains of 3.47 points at 10% retained KV and 3.66 points at 20% retained KV. The method preserves the compressed baseline’s peak memory footprint and stable long-context decoding, although its residual branch adds moderate throughput cost.
Original abstract
KV cache compression is essential for efficient long-context inference. Existing eviction methods permanently discard unselected tokens and consequently remove their aggregate contribution to attention. Merging-based alternatives preserve more information but can perturb retained keys and values that should remain exact. We observe that the information omitted by cache eviction can be formulated as residual statistics in both the numerator and denominator of softmax attention. Based on this observation, we propose ResKV, which divides a fixed KV budget into an exact main cache and a compact residual cache that reconstructs the contribution of omitted tokens. ResKV lets main-cache tokens and residual entries participate in the same softmax normalization, so residual entries restore both attention numerator and denominator mass rather than acting as a post-hoc correction. A construction-time validation proxy determines residual allocation for each layer and KV head, while a decode-time dynamic gate adjusts residual contributions for individual queries. Comprehensive evaluations on LongBench and RULER, covering query-aware and query-agnostic settings, multiple backbones, cache budgets, and representative compression baselines, demonstrate broad improvements under the same retained KV budget while preserving the practical efficiency of compressed decoding, including peak memory usage and long-context decode throughput.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.