NTH

LongAttnComp: Cross-Family Context Compression for Long-Context Reasoning

AuthorsMengmeng Ji, Ravi Shanker Raju, Jonathan Lingjie Li, Chen Wu

June 29, 2026 2 min read
Watch on YouTube
The one-line take

LongAttnComp makes very long prompts cheaper to process by learning a smarter way to keep the most useful tokens, helping models reason over 100k+ context with less accuracy loss.

Key results

32000
Stage 1 training set

SQuAD + HotpotQA examples used to build the retrieval foundation

20000
Stage 2 training set

MuSiQue, 2WikiMultiHopQA, and replay data used for multi-hop extension

74.37
Code-Debug accuracy full context

DeepSeek-R1-0528 baseline on InfiniteBench Code-Debug

75.38
Code-Debug accuracy Stage 1

LongAttnComp Stage 1 on DeepSeek-R1-0528, exceeding full context

76.90
Code-Debug accuracy Stage 2 subq

Best reported Code-Debug result on DeepSeek-R1-0528

48.9
LongBench v2 accuracy

Stage 2 subq with budget-only selection on DeepSeek-V3.1

What the paper found

LongAttnComp, from SambaNova Systems, adapts AttnComp into a cross-family long-context compressor for 100k+ token reasoning by fine-tuning only a lightweight cross-attention scorer on top of frozen Llama-3.1-8B-Instruct and operating at token-chunk level instead of document level. The method adds a token-budget top-p selector, positional reordering, and a format-agnostic query parser, then trains in two stages: Stage 1 on 32,000 SQuAD and HotpotQA examples to build general retrieval, and Stage 2 on 20,000 MuSiQue, 2WikiMultiHopQA, and replay data to strengthen multi-hop reasoning. On InfiniteBench Code-Debug, Stage 1 reaches 75.38% on DeepSeek-R1-0528, surpassing full-context 74.37% and beating Speculative Prefill 62.44%; Stage 2 subq improves to 76.90%. The same compressor transfers without retraining across DeepSeek-R1-0528, DeepSeek-V3.1, MiniMax-M2.5, and GPT-OSS-120B, with gains of 7–31 points over Speculative Prefill. On LongBench v2, Stage 2 with budget-only selection and a 512-token parsed query climbs to 48.9% on DeepSeek-V3.1 and 49.7% on DeepSeek-R1-0528, recovering 7–12 points over Stage 1. The paper’s main technical takeaway is that task coverage is driven more by training-data composition than by architectural limits, and that inference settings such as chunk size and selection mode must match whether evidence is concentrated or distributed.

Original abstract

As real-world applications increasingly require processing inputs of 100k+ tokens, the gap between context length and inference efficiency has become a critical bottleneck. Context compression offers a way to reduce prefill costs while preserving task accuracy. However, existing training-free attention-based methods leave substantial gaps in demanding long-context tasks such as code reasoning. We present LongAttnComp, a long-context adaptation of AttnComp that fine-tunes a lightweight cross-attention scoring layer and introduces tokenlevel chunking, a token-budget top-p algorithm, positional reordering, and a formatagnostic query parser. We further design a two-stage fine-tuning recipe for the compressor: Stage 1 builds a general retrieval foundation from NIAH-style data, and Stage 2 extends it with multi-hop and reasoning data for broader long-context task coverage. On InfiniteBench Code-Debug, LongAttnComp matches or exceeds full-context accuracy, substantially outperforms training-free baselines, and transfers across four target models from three families. On LongBench v2, the two-stage recipe largely closes the Stage 1 gap on multi-document reasoning while preserving Code-Debug performance.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis