CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM
AuthorsYubo Li, Yidi Miao
Resources
CONF-KV makes long-context LLMs cheaper by using the model’s own confidence to decide how much memory to keep, helping preserve accuracy while cutting KV-cache use.
Key results
CONF-KV on WikiText-2 continuation at 2048 generated tokens
Fixed 512-token sliding-window baseline on GPT-2 at 2048 generated tokens
CONF-KV average retrieval accuracy up to 32K tokens
Fixed 512-token sliding-window baseline on Needle-in-a-Haystack
Fraction of full-KV task success retained by CONF-KV on 75 tasks
Pearson correlation between confidence score and KL shift after ablating the past 256 tokens on GPT-2
What the paper found
CONF-KV is a training-free KV-cache manager for long-horizon LLM inference that uses the target model’s own next-token distribution as a real-time confidence signal to decide when to keep a tight or loose cache budget, then ranks tokens by a mixture of EMA attention mass and recency while protecting the newest context. The system is paired with blockwise online-softmax attention, contiguous cache compaction, and mixed FP16/INT8 storage, and its pyramidal variant CONF-KV-L further allocates smaller budgets in deeper layers. Across GPT-2, Qwen-14B, OpenAI gpt-oss-20b, and Qwen-32B, the method stays close to a fixed 512-token sliding-window memory footprint while substantially improving quality: on GPT-2 at 2048 generated tokens it reaches 30.92 perplexity at 52.8 MB, and CONF-KV+INT8 reaches 31.26 perplexity at 38.7 MB, versus 34.37 for sliding-512. On Needle-in-a-Haystack up to 32K tokens, CONF-KV achieves 91.4% retrieval accuracy compared with 53.8% for sliding windows and 80.6% for H2O. On 75 VisualWebArena tasks, it preserves 95.3% of full-KV success while cutting peak memory by 2.8×. A mechanistic ablation shows confidence anticorrelates with KL shift after removing the past 256 tokens, with Pearson r = −0.77 on GPT-2, supporting the central assumption that low-confidence steps are the ones that need more retained context.
Original abstract
Long-horizon LLM inference turns the key--value (KV) cache into the dominant GPU memory consumer and makes per-token attention increasingly expensive. Many common eviction policies use static recency windows or historical attention, leaving unused a signal computed on every decoding step: the model's current uncertainty. We introduce CONF-KV, a KV-cache manager that converts the next-token distribution into a scalar confidence score and uses it to choose the per-step cache budget, retaining more context when the model is uncertain and pruning aggressively when it is confident. Within each budget, tokens are ranked by a composite of accumulated attention mass and recency, while a protected recent window preserves local coherence. We combine the policy with blockwise online-softmax attention, mixed FP16/INT8 storage, and a pyramidal per-layer budget variant. Across four model families and generated lengths up to 4K, CONF-KV stays near the footprint of a fixed 512-token sliding window while remaining within 1.5--2.1 perplexity points of full KV. On Needle-in-a-Haystack up to 32K tokens, CONF-KV reaches 91.4% retrieval accuracy versus 53.8% for sliding windows and 80.6% for H2O; on 75 VisualWebArena tasks it retains 95.3% of full-KV success at 2.8 times lower peak memory.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.