RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction
AuthorsChangwoo Baek, Seungjun Shin, Kyeongbo Kong
Resources
RestoreKV makes aggressively compressed LLM KV caches behave more like full caches by adding a compact, context-specific learned restoration cache.
Key results
LoRA adapters and restore-token embeddings trained for Qwen3-4B
Score at a 5% KV budget before restoration
Score at a 5% KV budget on Qwen3-4B
RULER accuracy for RestoreKV+ at 16× compression
Maximum reported overhead for 32K-context restoration
What the paper found
RestoreKV addresses a central weakness of query-agnostic KV-cache eviction: selecting fewer original states can severely damage long-context behavior when compression becomes aggressive. The method keeps the base evictor, reserves 8 slots within the same total KV budget, and uses 8 learned restore tokens with LoRA to attend once to the full pre-eviction cache, producing a context-conditioned complement; the adapters are then disabled for all queries and decoding. Training uses symmetric-KL self-distillation from the frozen full-cache model, updating only 0.4% of the Qwen3-4B backbone without task-specific tuning. On Qwen3-4B, RestoreKV raises KVzip’s RULER-4K score from 38.2 to 73.2 at a 5% budget, and RestoreKV+ reaches 86.4 RULER accuracy at 16× compression on the KVPress Benchmark. The approach generalizes across Qwen3-0.6B, Qwen3-4B, Qwen3-8B, and Meta’s Llama-3.1-8B-Instruct, as well as five eviction methods and four long-context benchmarks. For 32K-context evaluation on an NVIDIA RTX PRO 6000, restoration adds only 0.5% of total compression time and no query-time KV-memory or decoding cost, supporting the paper’s claim that attention-side adaptation can recover information lost by eviction rather than merely improving token selection.
Original abstract
Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future queries, but performance can collapse under tight budgets. Existing methods primarily improve which original KV pairs are retained. We introduce RestoreKV, which complements this selection-based formulation with learned restoration under the same total KV budget. Our key insight is that, although the information lost through eviction is context-specific, the mechanism for generating its compact complement can be shared across contexts. After context prefill, a few restore tokens attend to the full KV cache in a single LoRA-adapted pass, generating a compact, context-conditioned restore cache. The base importance scorer and eviction rule remain unchanged, and the adapters are disabled for all subsequent queries and decoding. RestoreKV is trained through parameter-efficient self-distillation from the frozen full-cache model, optimizing only $0.4\%$ of the parameters and requiring no task-specific tuning. Across four backbones and four long-context benchmarks, RestoreKV substantially reduces compression-induced degradation. On Qwen3-4B, it improves 59 of 60 paired, budget-matched settings across five base eviction methods; at a $5\%$ budget, it raises KVzip from $38.2$ to $73.2$ on RULER-4K. Applied to KVzip+, RestoreKV reaches $86.4$ RULER accuracy at $16\times$ compression on the KVPress Benchmark, while adding less than $0.5\%$ one-time cache-construction overhead in a 32K-context evaluation. Our project page is available at https://paper.pnu-cvsp.com/RestoreKV/
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.