Still: Amortized KV Cache Compaction in a Single Forward Pass
AuthorsCharles O'Neill, Alex Sandomirsky, Harry Partridge, Mudith Jayasekara, Max Kirkby
Resources
Still makes long-context language models more practical by compressing their key-value cache in one fast forward pass, preserving performance even at extreme compression ratios.
Key results
Approximate size of the Qwen3-4B Still compactor at t = 128
Maximum compression ratio evaluated in the main sweep
Maximum context length evaluated in the main sweep
Largest reported Still over KV-Distill gain on matched-training RULER
Number of matched-training RULER cells where Still exceeds KV-Distill
Pairwise judge wins for Still against KV-Distill on LongBench v1 summarization
What the paper found
Still, developed by Baseten, is a Perceiver-based KV cache compactor that amortizes synthesis into a single forward pass instead of per-context optimization. For frozen Qwen3 and Gemma models, it inverse-rotates cached keys into a position-free frame, applies per-layer latent cross-attention, and writes compact keys and values that can be reused across inference, including iterative long-horizon compaction. On Qwen3-4B, the canonical compactor uses 256 latent slots and about 50M parameters, roughly 1% of the base model, and the paper evaluates compression ratios from 8× to 200× over 8k to 128k contexts. Still stays on the favorable side of the speed–quality frontier, and on the long-context RULER grid it exceeds KV-Distill by 8–22 accuracy points in 16 of 18 matched-training cells. In free-form summarization, it preserves 74–95% of the full-context gain on HELMET multi_lexsum from 8k to 64k and wins 300 of 500 pairwise judgments against KV-Distill on LongBench v1 GovReport and QMSum at 16k. The paper’s main technical message is that learned synthesis, not token selection, makes compact KV state expressive enough to survive extreme compression while remaining practical for repeated inference-time use.
Original abstract
The KV cache is the memory bottleneck of long-horizon language model deployment. Practically, a deployable compactor must be lightweight enough to call during inference, expressive enough to preserve context under constraint, and reusable across a trajectory. Existing compaction methods satisfy only part of this requirement: selection methods are lightweight but subset-bound, while synthesis methods are expressive but rely on per-context optimization. Here we introduce Still, a small per-layer Perceiver trained once against a frozen base model that produces compact keys and values in a single forward pass. On Qwen and Gemma models, Still occupies the favorable side of the speed--quality frontier across compression ratios from $8\times$ to $200\times$ and context lengths from $8$k to $128$k. On the long-context RULER grid, Still exceeds the strongest baseline by 8--22 points. The same compact cache also supports free-form summarization, preserving most of the full-context gain on HELMET and winning a pairwise LongBench summarization comparison against KV-Distill. Because compaction is a forward pass, Still can be applied iteratively, entering a long-horizon regime unavailable to per-context methods. We show that amortization makes long-context cache compaction tractable, and synthesis makes its compact state useful at extreme compression.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.