NTH

RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction

AuthorsChangwoo Baek, Seungjun Shin, Kyeongbo Kong

August 8, 2026 2 min read
Watch on YouTube
The one-line take

RestoreKV makes aggressively compressed LLM KV caches behave more like full caches by adding a compact, context-specific learned restoration cache.

Key results

0.4%
Trainable parameter fraction

LoRA adapters and restore-token embeddings trained for Qwen3-4B

38.2
KVzip RULER-4K baseline

Score at a 5% KV budget before restoration

73.2
RestoreKV RULER-4K score

Score at a 5% KV budget on Qwen3-4B

86.4
KVPress Benchmark accuracy

RULER accuracy for RestoreKV+ at 16× compression

0.5%
Compression-time overhead

Maximum reported overhead for 32K-context restoration

What the paper found

RestoreKV addresses a central weakness of query-agnostic KV-cache eviction: selecting fewer original states can severely damage long-context behavior when compression becomes aggressive. The method keeps the base evictor, reserves 8 slots within the same total KV budget, and uses 8 learned restore tokens with LoRA to attend once to the full pre-eviction cache, producing a context-conditioned complement; the adapters are then disabled for all queries and decoding. Training uses symmetric-KL self-distillation from the frozen full-cache model, updating only 0.4% of the Qwen3-4B backbone without task-specific tuning. On Qwen3-4B, RestoreKV raises KVzip’s RULER-4K score from 38.2 to 73.2 at a 5% budget, and RestoreKV+ reaches 86.4 RULER accuracy at 16× compression on the KVPress Benchmark. The approach generalizes across Qwen3-0.6B, Qwen3-4B, Qwen3-8B, and Meta’s Llama-3.1-8B-Instruct, as well as five eviction methods and four long-context benchmarks. For 32K-context evaluation on an NVIDIA RTX PRO 6000, restoration adds only 0.5% of total compression time and no query-time KV-memory or decoding cost, supporting the paper’s claim that attention-side adaptation can recover information lost by eviction rather than merely improving token selection.

Original abstract

Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future queries, but performance can collapse under tight budgets. Existing methods primarily improve which original KV pairs are retained. We introduce RestoreKV, which complements this selection-based formulation with learned restoration under the same total KV budget. Our key insight is that, although the information lost through eviction is context-specific, the mechanism for generating its compact complement can be shared across contexts. After context prefill, a few restore tokens attend to the full KV cache in a single LoRA-adapted pass, generating a compact, context-conditioned restore cache. The base importance scorer and eviction rule remain unchanged, and the adapters are disabled for all subsequent queries and decoding. RestoreKV is trained through parameter-efficient self-distillation from the frozen full-cache model, optimizing only $0.4\%$ of the parameters and requiring no task-specific tuning. Across four backbones and four long-context benchmarks, RestoreKV substantially reduces compression-induced degradation. On Qwen3-4B, it improves 59 of 60 paired, budget-matched settings across five base eviction methods; at a $5\%$ budget, it raises KVzip from $38.2$ to $73.2$ on RULER-4K. Applied to KVzip+, RestoreKV reaches $86.4$ RULER accuracy at $16\times$ compression on the KVPress Benchmark, while adding less than $0.5\%$ one-time cache-construction overhead in a 32K-context evaluation. Our project page is available at https://paper.pnu-cvsp.com/RestoreKV/

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis