FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
AuthorsYan Wang, Qifan Zhang, Jiachen Yu, Tian Liang, Dongyang Ma, Xiang Hu, Zibo Lin, Chunyang Li, Zhichao Wang, Jia Li, Yujiu Yang, Haitao Mi, Dong Yu
Resources
This paper introduces a new way to serve ultra-long-context LLMs by predicting which past tokens actually matter, cutting KV-cache memory drastically while keeping accuracy high.
Key results
Average physical KV cache retained by FM-DS-V4 versus the full-context baseline across LongBench-v2, LongMemEval, and RULER
Average reduction in GPU memory usage from the baseline across the primary benchmarks
FM-DS-V4 average accuracy across LongBench-v2, LongMemEval, and RULER
FM-DS-V4 accuracy at 493K context on LongBench-v2-L
FM-DS-V4 GPU memory overhead in GB at 493K context on LongBench-v2-L
Scale of the lookahead dataset used to train the Memory Indexer
What the paper found
FlashMemory-DeepSeek-V4, developed on the DeepSeek-V4-Flash stack with Tencent-affiliated researchers, proposes Lookahead Sparse Attention (LSA), a Neural Memory Indexer that predicts which historical KV chunks will matter in the next decoding window and loads only those chunks into GPU memory. Instead of scanning the full cache, the indexer runs every 64 steps, uses a Sigmoid-thresholded dual-encoder trained separately from the backbone, and keeps DeepSeek-V4’s 128:1 Heavily Compressed Attention layers intact. The training pipeline builds golden labels from 10,000 long documents spanning 16K to 512K tokens using cross-layer majority voting over 21 CSA layers, then optimizes only the query-side projections with a lightweight focal-loss retrieval objective. Across LongBench-v2, LongMemEval, and RULER, FM-DS-V4 cuts average KV cache footprint to 13.5% of the full-context baseline, an 86.5% reduction, while improving average accuracy from 76.9% to 77.5%. At 493K context on LongBench-v2-L, it reaches 70.0% accuracy versus 68.1% for DS-V4-Flash while using 0.18 GB instead of 1.80 GB. The report also shows a 500K-scale setting where memory reduction exceeds 90%, but notes failure modes on MRCR and a length generalization ceiling beyond roughly 2× training context.
Original abstract
Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In this report, we propose Lookahead Sparse Attention (LSA), a novel inference paradigm powered by a Neural Memory Indexer built upon the DeepSeek-V4 architecture. Rather than passively attending to all historical tokens, LSA proactively predicts future context demands and preserves only the query-critical KV chunks in the GPU memory. Crucially, we instantiate this architecture via a backbone-free decoupled training strategy. By formulating the indexer as a standard dual-encoder architecture, we train it independently using standard retrieval training frameworks without ever loading the massive backbone model into GPU memory. We demonstrate that this "less is more" paradigm significantly maximizes serving efficiency while acting as an effective attention denoiser in tasks that rely on long-term global memory. Across primary long-context evaluation suites (e.g., LongBench-v2, LongMemEval, and RULER), FM-DS-V4 compresses the average physical KV cache footprint down to merely 13.5% of the full-context baseline, while consistently preserving or slightly elevating downstream accuracy (+0.6% absolute margin on average). Crucially, at extreme 500K scales, FlashMemory suppresses the physical KV cache overhead by over 90% without destabilizing the backbone's core reasoning capacities.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.