NTH

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM

AuthorsYubo Li, Yidi Miao

June 17, 2026 2 min read
Watch on YouTube
The one-line take

CONF-KV makes long-context LLMs cheaper by using the model’s own confidence to decide how much memory to keep, helping preserve accuracy while cutting KV-cache use.

Key results

30.92
GPT-2 perplexity

CONF-KV on WikiText-2 continuation at 2048 generated tokens

34.37
Sliding-512 perplexity

Fixed 512-token sliding-window baseline on GPT-2 at 2048 generated tokens

91.4%
Needle-in-a-Haystack accuracy

CONF-KV average retrieval accuracy up to 32K tokens

53.8%
Sliding-512 NIAH accuracy

Fixed 512-token sliding-window baseline on Needle-in-a-Haystack

95.3%
VisualWebArena success retention

Fraction of full-KV task success retained by CONF-KV on 75 tasks

-0.77
Confidence-KL correlation

Pearson correlation between confidence score and KL shift after ablating the past 256 tokens on GPT-2

What the paper found

CONF-KV is a training-free KV-cache manager for long-horizon LLM inference that uses the target model’s own next-token distribution as a real-time confidence signal to decide when to keep a tight or loose cache budget, then ranks tokens by a mixture of EMA attention mass and recency while protecting the newest context. The system is paired with blockwise online-softmax attention, contiguous cache compaction, and mixed FP16/INT8 storage, and its pyramidal variant CONF-KV-L further allocates smaller budgets in deeper layers. Across GPT-2, Qwen-14B, OpenAI gpt-oss-20b, and Qwen-32B, the method stays close to a fixed 512-token sliding-window memory footprint while substantially improving quality: on GPT-2 at 2048 generated tokens it reaches 30.92 perplexity at 52.8 MB, and CONF-KV+INT8 reaches 31.26 perplexity at 38.7 MB, versus 34.37 for sliding-512. On Needle-in-a-Haystack up to 32K tokens, CONF-KV achieves 91.4% retrieval accuracy compared with 53.8% for sliding windows and 80.6% for H2O. On 75 VisualWebArena tasks, it preserves 95.3% of full-KV success while cutting peak memory by 2.8×. A mechanistic ablation shows confidence anticorrelates with KL shift after removing the past 256 tokens, with Pearson r = −0.77 on GPT-2, supporting the central assumption that low-confidence steps are the ones that need more retained context.

Original abstract

Long-horizon LLM inference turns the key--value (KV) cache into the dominant GPU memory consumer and makes per-token attention increasingly expensive. Many common eviction policies use static recency windows or historical attention, leaving unused a signal computed on every decoding step: the model's current uncertainty. We introduce CONF-KV, a KV-cache manager that converts the next-token distribution into a scalar confidence score and uses it to choose the per-step cache budget, retaining more context when the model is uncertain and pruning aggressively when it is confident. Within each budget, tokens are ranked by a composite of accumulated attention mass and recency, while a protected recent window preserves local coherence. We combine the policy with blockwise online-softmax attention, mixed FP16/INT8 storage, and a pyramidal per-layer budget variant. Across four model families and generated lengths up to 4K, CONF-KV stays near the footprint of a fixed 512-token sliding window while remaining within 1.5--2.1 perplexity points of full KV. On Needle-in-a-Haystack up to 32K tokens, CONF-KV reaches 91.4% retrieval accuracy versus 53.8% for sliding windows and 80.6% for H2O; on 75 VisualWebArena tasks it retains 95.3% of full-KV success at 2.8 times lower peak memory.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis