NTH

IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference

AuthorsXintong Yang, Hao Gu, Binxing Xu, Lujun Li, Bei Liu, Jiacheng Liu, Qiyuan Zhu, Sirui Han, Yike Guo

May 30, 2026 2 min read
Watch on YouTube
The one-line take

IndexMem makes long-context LLMs more efficient by learning what to keep in memory and compressing what gets evicted so the model can still recall important information later.

Key results

25
RULER gain

IndexMem reports gains up to 25 points on RULER across Qwen, Mistral, and Llama under aggressive eviction.

55.26
LongBench average

At 50% compression, IndexMem achieves a LongBench average score of 55.26.

42.35
PyramidKV LongBench average

At 50% compression on LongBench, PyramidKV scores 42.35 for comparison against IndexMem.

42.10
TOVA LongBench average

At 50% compression on LongBench, TOVA scores 42.10 for comparison against IndexMem.

19.92M
Indexer parameters

On Llama-3.1-8B-Instruct, the learnable indexer adds 19.92M parameters.

0.52M
Memory parameters

The latent memory module adds only 0.52M parameters beyond the indexer.

What the paper found

IndexMem targets the KV-cache memory wall in long-context LLM inference by replacing heuristic eviction with a learnable indexer and adding a latent memory that preserves evicted information instead of deleting it permanently. The indexer is a lightweight MQA-style module trained by streaming KL distillation to match the backbone’s attention-derived token-importance distribution, using QK-Norm features and max-over-queries aggregation; during decoding it can reuse cached keys and even compress before prefill ends. The memory module stores evicted tokens in a fixed-size fast-weight state updated online with decay, then returns a gated residual readout, so the final output is attention over retained KV plus compensation from latent memory rather than brittle “memory-as-tokens” softmax insertion. On Qwen3-8B, Mistral-7B-v0.3, and Llama-3.1-8B-Instruct, evaluated with NVIDIA’s KVPRESS on RULER, Needle-in-a-Haystack, and LongBench, IndexMem is consistently the strongest eviction method: it nearly matches full-cache accuracy under 10–25% compression and holds up much better at 50–90% compression, with gains up to 25 points on RULER. The memory residual reduces catastrophic retrieval failures, improves LongBench average scores to 55.26 at 50% compression versus 42.35 for PyramidKV and 42.10 for TOVA, and adds only 0.52M parameters beyond the 19.92M-parameter indexer.

Original abstract

Large Language Models (LLMs) are increasingly expected to operate over long contexts, yet standard softmax attention incurs a KV cache that grows linearly with sequence length, quickly becoming the bottleneck for long context inference. A practical remedy is to evict less important KV entries; however, existing eviction policies are largely heuristic and struggle to capture the rich, input-dependent distribution of token importance. In this work, we introduce a learnable indexer that predicts KV importance, enabling more accurate retention of critical tokens. Meanwhile, naively evicting tokens permanently discards their information, leading to irreversible forgetting and degraded retrieval over long ranges. To address this, we propose a lightweight latent memory module that compresses evicted tokens into a compact, online-updated state and provides residual readouts to compensate for the attention contributions lost through KV eviction. Collectively, our method enables accurate long-context inference under a bounded KV budget, delivering consistent improvements on RULER (4K/16K) across Qwen, Mistral, and Llama models (up to 25 points under aggressive eviction), markedly more stable Needle-in-a-Haystack retrieval, and superior LongBench scores and compression curves compared to existing eviction policies.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis