NTH

End-to-End Context Compression at Scale

AuthorsAng Li, Sean McLeish, Haozhe Chen, Nimit Kalra, Zaiqian Chen, Artem Gazizov, Venkata Anoop Suhas Kumar Morisetty, Bhavya Kailkhura, Harshitha Menon, Zhuang Liu, Brian R. Bartoldson, Tom Goldstein, Sanae Lotfi, Micah Goldblum, Pavel Izmailov

June 9, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces latent context models that compress long prompts into compact representations, making long-context language models faster and more memory-efficient without sacrificing as much quality.

Key results

0.6B
encoder size

encoder used in the main LCLM family

4B
decoder size

decoder used in the main LCLM family

350B
training tokens

joint encoder-decoder pretraining scale

75.06
RULER 4k

16x-compressed LCLM accuracy

39.08
LongBench English

16x-compressed LCLM average score

93.99
NIAH agent gain

RULER needle-in-a-haystack accuracy with expand-on-demand agent at 4k context

What the paper found

End-to-End Context Compression at Scale, from NYU, Harvard, Princeton, Columbia, Meta FAIR, and Lawrence Livermore National Laboratory, revisits encoder-decoder soft-token compression as a production-friendly alternative to KV-cache eviction. The paper introduces Latent Context Language Models, or LCLMs, built by jointly training a 0.6B encoder with a 4B decoder on more than 350B tokens, and it shows that a carefully staged recipe matters: adapter warmup, encoder training, end-to-end continual pre-training, and supervised fine-tuning with interleaved compressed and uncompressed spans. A controlled architecture search finds that causal encoder masking, a 1024-token encoder window, and a lightweight MLP adapter outperform bidirectional masking, smaller windows, and attention-based adapters. On long-context benchmarks, LCLMs define a new speed-memory-accuracy frontier: at 16× compression they reach 75.06 on RULER 4k, 70.23 on RULER 8k, 65.91 on RULER 16k, 39.08 on LongBench English, 31.74 on LongBench Chinese, 67.50 on LongHealth5, and 81.05 on GSM8K, while substantially reducing time-to-first-token and peak GPU memory versus SnapKV, KVzip, FastKVzip, Expected Attention, and Attention Matching. The paper also shows that the same compressor can support an agentic expand-on-demand loop over 512-token chunks, boosting RULER needle-in-a-haystack retrieval from 76.89 to 93.99 at 4k context and from 70.18 to 89.91 at 16k context, demonstrating that compressed latent memory can preserve global coverage while recovering exact details only when needed.

Original abstract

Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length. Recent techniques to compress the KV cache fall short: they either degrade model quality substantially or require considerable time and compute to compress a single long prompt. Furthermore, many methods require the input to fit within the target model's context window, and are generally incompatible with modern production inference engines. Encoder-decoder compressors, which map a long token sequence to a shorter sequence of latent embeddings consumed by a decoder, are an appealing alternative in principle. However, existing approaches are not competitive with KV cache compression on the accuracy-efficiency frontier. In this work, we revisit encoder-decoder compression and close this gap. We first perform an architecture search, pre-training many variants from scratch to determine how best to design and train encoder-decoder compressors. Guided by our findings, we continually pre-train a family of 0.6B-encoder, 4B-decoder models on over 350B tokens each, at compression ratios of 1:4, 1:8, and 1:16. We introduce Latent Context Language Models (LCLMs), a family of compressors that improve the Pareto frontier across general-task performance, compression speed, and peak memory usage. We demonstrate that LCLMs serve as efficient backbones for long-horizon agents, letting the agent skim through a compressed long context and adaptively expand relevant segments on demand.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis