NTH

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

AuthorsYan Wang, Qifan Zhang, Jiachen Yu, Tian Liang, Dongyang Ma, Xiang Hu, Zibo Lin, Chunyang Li, Zhichao Wang, Jia Li, Yujiu Yang, Haitao Mi, Dong Yu

June 21, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces a new way to serve ultra-long-context LLMs by predicting which past tokens actually matter, cutting KV-cache memory drastically while keeping accuracy high.

Key results

13.5%
avg KV cache footprint

Average physical KV cache retained by FM-DS-V4 versus the full-context baseline across LongBench-v2, LongMemEval, and RULER

86.5%
avg memory reduction

Average reduction in GPU memory usage from the baseline across the primary benchmarks

77.5%
avg accuracy

FM-DS-V4 average accuracy across LongBench-v2, LongMemEval, and RULER

70.0%
LongBench-v2-L accuracy

FM-DS-V4 accuracy at 493K context on LongBench-v2-L

0.18
LongBench-v2-L memory

FM-DS-V4 GPU memory overhead in GB at 493K context on LongBench-v2-L

10000
training documents

Scale of the lookahead dataset used to train the Memory Indexer

What the paper found

FlashMemory-DeepSeek-V4, developed on the DeepSeek-V4-Flash stack with Tencent-affiliated researchers, proposes Lookahead Sparse Attention (LSA), a Neural Memory Indexer that predicts which historical KV chunks will matter in the next decoding window and loads only those chunks into GPU memory. Instead of scanning the full cache, the indexer runs every 64 steps, uses a Sigmoid-thresholded dual-encoder trained separately from the backbone, and keeps DeepSeek-V4’s 128:1 Heavily Compressed Attention layers intact. The training pipeline builds golden labels from 10,000 long documents spanning 16K to 512K tokens using cross-layer majority voting over 21 CSA layers, then optimizes only the query-side projections with a lightweight focal-loss retrieval objective. Across LongBench-v2, LongMemEval, and RULER, FM-DS-V4 cuts average KV cache footprint to 13.5% of the full-context baseline, an 86.5% reduction, while improving average accuracy from 76.9% to 77.5%. At 493K context on LongBench-v2-L, it reaches 70.0% accuracy versus 68.1% for DS-V4-Flash while using 0.18 GB instead of 1.80 GB. The report also shows a 500K-scale setting where memory reduction exceeds 90%, but notes failure modes on MRCR and a length generalization ceiling beyond roughly 2× training context.

Original abstract

Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In this report, we propose Lookahead Sparse Attention (LSA), a novel inference paradigm powered by a Neural Memory Indexer built upon the DeepSeek-V4 architecture. Rather than passively attending to all historical tokens, LSA proactively predicts future context demands and preserves only the query-critical KV chunks in the GPU memory. Crucially, we instantiate this architecture via a backbone-free decoupled training strategy. By formulating the indexer as a standard dual-encoder architecture, we train it independently using standard retrieval training frameworks without ever loading the massive backbone model into GPU memory. We demonstrate that this "less is more" paradigm significantly maximizes serving efficiency while acting as an effective attention denoiser in tasks that rely on long-term global memory. Across primary long-context evaluation suites (e.g., LongBench-v2, LongMemEval, and RULER), FM-DS-V4 compresses the average physical KV cache footprint down to merely 13.5% of the full-context baseline, while consistently preserving or slightly elevating downstream accuracy (+0.6% absolute margin on average). Crucially, at extreme 500K scales, FlashMemory suppresses the physical KV cache overhead by over 90% without destabilizing the backbone's core reasoning capacities.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis