IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
AuthorsXintong Yang, Hao Gu, Binxing Xu, Lujun Li, Bei Liu, Jiacheng Liu, Qiyuan Zhu, Sirui Han, Yike Guo
Resources
IndexMem makes long-context LLMs more efficient by learning what to keep in memory and compressing what gets evicted so the model can still recall important information later.
Key results
IndexMem reports gains up to 25 points on RULER across Qwen, Mistral, and Llama under aggressive eviction.
At 50% compression, IndexMem achieves a LongBench average score of 55.26.
At 50% compression on LongBench, PyramidKV scores 42.35 for comparison against IndexMem.
At 50% compression on LongBench, TOVA scores 42.10 for comparison against IndexMem.
On Llama-3.1-8B-Instruct, the learnable indexer adds 19.92M parameters.
The latent memory module adds only 0.52M parameters beyond the indexer.
What the paper found
IndexMem targets the KV-cache memory wall in long-context LLM inference by replacing heuristic eviction with a learnable indexer and adding a latent memory that preserves evicted information instead of deleting it permanently. The indexer is a lightweight MQA-style module trained by streaming KL distillation to match the backbone’s attention-derived token-importance distribution, using QK-Norm features and max-over-queries aggregation; during decoding it can reuse cached keys and even compress before prefill ends. The memory module stores evicted tokens in a fixed-size fast-weight state updated online with decay, then returns a gated residual readout, so the final output is attention over retained KV plus compensation from latent memory rather than brittle “memory-as-tokens” softmax insertion. On Qwen3-8B, Mistral-7B-v0.3, and Llama-3.1-8B-Instruct, evaluated with NVIDIA’s KVPRESS on RULER, Needle-in-a-Haystack, and LongBench, IndexMem is consistently the strongest eviction method: it nearly matches full-cache accuracy under 10–25% compression and holds up much better at 50–90% compression, with gains up to 25 points on RULER. The memory residual reduces catastrophic retrieval failures, improves LongBench average scores to 55.26 at 50% compression versus 42.35 for PyramidKV and 42.10 for TOVA, and adds only 0.52M parameters beyond the 19.92M-parameter indexer.
Original abstract
Large Language Models (LLMs) are increasingly expected to operate over long contexts, yet standard softmax attention incurs a KV cache that grows linearly with sequence length, quickly becoming the bottleneck for long context inference. A practical remedy is to evict less important KV entries; however, existing eviction policies are largely heuristic and struggle to capture the rich, input-dependent distribution of token importance. In this work, we introduce a learnable indexer that predicts KV importance, enabling more accurate retention of critical tokens. Meanwhile, naively evicting tokens permanently discards their information, leading to irreversible forgetting and degraded retrieval over long ranges. To address this, we propose a lightweight latent memory module that compresses evicted tokens into a compact, online-updated state and provides residual readouts to compensate for the attention contributions lost through KV eviction. Collectively, our method enables accurate long-context inference under a bounded KV budget, delivering consistent improvements on RULER (4K/16K) across Qwen, Mistral, and Llama models (up to 25 points under aggressive eviction), markedly more stable Needle-in-a-Haystack retrieval, and superior LongBench scores and compression curves compared to existing eviction policies.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.