NTH

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

AuthorsDeepSeek-AI, :, Anyi Xu, B. Li, Bangcai Lin, Bing Xue, BingCheng Xian, Bingzheng Xu, Bochao Wu, Bowei Zhang, Boyi Deng, C. C. Yu, Chao Jin, Chaofan Lin, Chen Dong, Chenbing Wang, Chenfan Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chengyuan Zhang, Chenhao Xu, Chenqi Zhao, Chenze Shao, Chuhao Wang, Chuqi Zhang, Damai Dai, Dejian Yang, Deli Chen, Di Huang, Di Wu, Donghao Li, Erhang Li, Eric Fu, F. Zhou, Fangwei Zhou, Fangyun Lin, Fangzhou Yuan, Feiyu Xia, Fucong Dai, Guangbo Hao, Guanglin Li, Guanting Chen, Guoai Cao, Guofan Fan, Guolai Meng, Guowei Li, Haichuan Zhang, Haiyang Ma, Haiyang Shen, Han Li, Han Yu, Han Zhang, Hangyuan Deng, Hanwei Xu, Hanxiang Xu, Hanxun Zhong, Hao Guo, Hao Jiang, Hao Li, Hao Qin, Haodong Wen, Haofen Liang, Haofeng Huang, Haohua Liu, Haoling Zhang, Haoming Luo, Haoran Yang, Haotian Xu, Haotian Yuan, Haoting Huang, Haowen Luo, Haoyang Cai, Haoyu Chen, Haozhe Ji, Hengran Zhang, Hengrui Wang, Hengxu Wu, Honghui Ding, Hongxuan Tang, Huadong Wang, Huanqi ...

September 18, 2026 3 min read
Watch on YouTube
The one-line take

DeepSeek-V4.1-Flash targets million-token agents by dramatically compressing KV caches while maintaining strong multimodal performance.

Key results

552B
Backbone parameters

Total backbone parameter count of DeepSeek V4.1 Flash.

890
Global KV footprint

Global KV cache storage in bytes per token after CSA2 and FP4 caching.

45T
Pretraining corpus

Multimodal tokens used for pretraining.

74.2
DeepSWE v1.1

Resolved-task performance at maximum reasoning effort.

90.6
Terminal-Bench 2.1

Pass@1 performance at maximum reasoning effort.

76.3%
Reasoning-effort gain

Average Pass@1 across eight reasoning benchmarks at effort 100.

What the paper found

DeepSeek V4.1 Flash targets the main deployment bottleneck for long-horizon agents: prefill computation and KV-cache storage. It is a multimodal Mixture-of-Experts model with 552B backbone parameters, support for up to one million tokens, and a Causal Encoder-Decoder architecture that activates 8B parameters per token during prefill and 16B during decode. Its main contribution is aggressive cache compression: Compressed Sparse Attention 2, or CSA2, reuses global KV states, indexer keys, and Top-K selections across layers, while FP4 quantization reduces the global cache to 890 bytes per token—roughly one-quarter of DeepSeek-V4-Flash’s footprint. SWA Bounded Replay avoids persistently storing sliding-window attention states, reducing persistent KV storage to roughly one-eighth of the predecessor’s requirement with negligible quality loss. The model was pretrained on 45T multimodal tokens and combines Single-Pass mHC, Engram conditional memory, DSpark speculative decoding, and hierarchical sparse indexing. Despite its smaller active footprint, it reaches 74.2% on DeepSWE v1.1 and 90.6% on Terminal-Bench 2.1, competitive with systems such as Anthropic’s Claude and OpenAI’s GPT-5.6. A controllable reasoning-effort scalar also raises eight-benchmark average Pass@1 from 67.1% at effort 25 to 76.3% at effort 100, at approximately 2.5× the output-token cost, giving operators an explicit quality-versus-latency control.

Original abstract

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the primary bottleneck to further lowering deployment costs. To address this challenge, we introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. With its Causal Encoder-Decoder (CED) architecture, the model activates 16B parameters per token during decode but only 8B parameters during prefill, substantially improving cost efficiency for agentic workloads. To push the limits of KV cache compression, DeepSeek-V4.1-Flash combines cross-layer KV cache reuse in Compressed Sparse Attention 2 (CSA2) with FP4 KV caching. These designs reduce its global KV cache footprint (always in HBM) to 890 bytes per token, roughly 1/4 of the corresponding footprint of DeepSeek-V4-Flash. Further, through a dedicated deployment optimization known as SWA Bounded Replay, DeepSeek-V4.1-Flash reduces its persistent KV cache footprint (always on SSD or in host memory) to roughly 1/8 of that of DeepSeek-V4-Flash. Despite its much smaller KV cache footprint, the model delivers substantially better performance than the baseline. In addition, we streamline the DeepSeek-V4 architecture and introduce several efficient architectural extensions. We pretrain DeepSeek-V4.1-Flash on a multimodal corpus comprising 45T tokens and conduct comprehensive post-training, yielding strong performance across diverse text-based and multimodal agentic scenarios. Model checkpoints are available at https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis