End-to-End Context Compression at Scale
AuthorsAng Li, Sean McLeish, Haozhe Chen, Nimit Kalra, Zaiqian Chen, Artem Gazizov, Venkata Anoop Suhas Kumar Morisetty, Bhavya Kailkhura, Harshitha Menon, Zhuang Liu, Brian R. Bartoldson, Tom Goldstein, Sanae Lotfi, Micah Goldblum, Pavel Izmailov
Resources
This paper introduces latent context models that compress long prompts into compact representations, making long-context language models faster and more memory-efficient without sacrificing as much quality.
Key results
encoder used in the main LCLM family
decoder used in the main LCLM family
joint encoder-decoder pretraining scale
16x-compressed LCLM accuracy
16x-compressed LCLM average score
RULER needle-in-a-haystack accuracy with expand-on-demand agent at 4k context
What the paper found
End-to-End Context Compression at Scale, from NYU, Harvard, Princeton, Columbia, Meta FAIR, and Lawrence Livermore National Laboratory, revisits encoder-decoder soft-token compression as a production-friendly alternative to KV-cache eviction. The paper introduces Latent Context Language Models, or LCLMs, built by jointly training a 0.6B encoder with a 4B decoder on more than 350B tokens, and it shows that a carefully staged recipe matters: adapter warmup, encoder training, end-to-end continual pre-training, and supervised fine-tuning with interleaved compressed and uncompressed spans. A controlled architecture search finds that causal encoder masking, a 1024-token encoder window, and a lightweight MLP adapter outperform bidirectional masking, smaller windows, and attention-based adapters. On long-context benchmarks, LCLMs define a new speed-memory-accuracy frontier: at 16× compression they reach 75.06 on RULER 4k, 70.23 on RULER 8k, 65.91 on RULER 16k, 39.08 on LongBench English, 31.74 on LongBench Chinese, 67.50 on LongHealth5, and 81.05 on GSM8K, while substantially reducing time-to-first-token and peak GPU memory versus SnapKV, KVzip, FastKVzip, Expected Attention, and Attention Matching. The paper also shows that the same compressor can support an agentic expand-on-demand loop over 512-token chunks, boosting RULER needle-in-a-haystack retrieval from 76.89 to 93.99 at 4k context and from 70.18 to 89.91 at 16k context, demonstrating that compressed latent memory can preserve global coverage while recovering exact details only when needed.
Original abstract
Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length. Recent techniques to compress the KV cache fall short: they either degrade model quality substantially or require considerable time and compute to compress a single long prompt. Furthermore, many methods require the input to fit within the target model's context window, and are generally incompatible with modern production inference engines. Encoder-decoder compressors, which map a long token sequence to a shorter sequence of latent embeddings consumed by a decoder, are an appealing alternative in principle. However, existing approaches are not competitive with KV cache compression on the accuracy-efficiency frontier. In this work, we revisit encoder-decoder compression and close this gap. We first perform an architecture search, pre-training many variants from scratch to determine how best to design and train encoder-decoder compressors. Guided by our findings, we continually pre-train a family of 0.6B-encoder, 4B-decoder models on over 350B tokens each, at compression ratios of 1:4, 1:8, and 1:16. We introduce Latent Context Language Models (LCLMs), a family of compressors that improve the Pareto frontier across general-task performance, compression speed, and peak memory usage. We demonstrate that LCLMs serve as efficient backbones for long-horizon agents, letting the agent skim through a compressed long context and adaptively expand relevant segments on demand.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.