NTH

Still: Amortized KV Cache Compaction in a Single Forward Pass

AuthorsCharles O'Neill, Alex Sandomirsky, Harry Partridge, Mudith Jayasekara, Max Kirkby

June 17, 2026 2 min read
Watch on YouTube
The one-line take

Still makes long-context language models more practical by compressing their key-value cache in one fast forward pass, preserving performance even at extreme compression ratios.

Key results

50M
canonical params

Approximate size of the Qwen3-4B Still compactor at t = 128

200x
compression range

Maximum compression ratio evaluated in the main sweep

128k
context range

Maximum context length evaluated in the main sweep

22
RULER advantage

Largest reported Still over KV-Distill gain on matched-training RULER

16
RULER matched cells

Number of matched-training RULER cells where Still exceeds KV-Distill

300
LongBench pairwise wins

Pairwise judge wins for Still against KV-Distill on LongBench v1 summarization

What the paper found

Still, developed by Baseten, is a Perceiver-based KV cache compactor that amortizes synthesis into a single forward pass instead of per-context optimization. For frozen Qwen3 and Gemma models, it inverse-rotates cached keys into a position-free frame, applies per-layer latent cross-attention, and writes compact keys and values that can be reused across inference, including iterative long-horizon compaction. On Qwen3-4B, the canonical compactor uses 256 latent slots and about 50M parameters, roughly 1% of the base model, and the paper evaluates compression ratios from 8× to 200× over 8k to 128k contexts. Still stays on the favorable side of the speed–quality frontier, and on the long-context RULER grid it exceeds KV-Distill by 8–22 accuracy points in 16 of 18 matched-training cells. In free-form summarization, it preserves 74–95% of the full-context gain on HELMET multi_lexsum from 8k to 64k and wins 300 of 500 pairwise judgments against KV-Distill on LongBench v1 GovReport and QMSum at 16k. The paper’s main technical message is that learned synthesis, not token selection, makes compact KV state expressive enough to survive extreme compression while remaining practical for repeated inference-time use.

Original abstract

The KV cache is the memory bottleneck of long-horizon language model deployment. Practically, a deployable compactor must be lightweight enough to call during inference, expressive enough to preserve context under constraint, and reusable across a trajectory. Existing compaction methods satisfy only part of this requirement: selection methods are lightweight but subset-bound, while synthesis methods are expressive but rely on per-context optimization. Here we introduce Still, a small per-layer Perceiver trained once against a frozen base model that produces compact keys and values in a single forward pass. On Qwen and Gemma models, Still occupies the favorable side of the speed--quality frontier across compression ratios from $8\times$ to $200\times$ and context lengths from $8$k to $128$k. On the long-context RULER grid, Still exceeds the strongest baseline by 8--22 points. The same compact cache also supports free-form summarization, preserving most of the full-context gain on HELMET and winning a pairwise LongBench summarization comparison against KV-Distill. Because compaction is a forward pass, Still can be applied iteratively, entering a long-horizon regime unavailable to per-context methods. We show that amortization makes long-context cache compaction tractable, and synthesis makes its compact state useful at extreme compression.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis