MiniMax Sparse Attention
AuthorsXunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Pengyu Zhao
MiniMax Sparse Attention makes ultra-long-context LLMs much faster by combining blockwise sparse retrieval with GPU-friendly kernels, delivering major inference speedups at 1M-token scale.
Key results
main multimodal MoE model scale used for validation
total pretraining token budget
per-token attention FLOPs reduction at 1M context
wall-clock speedup on H800
wall-clock decoding speedup on H800
tokens attended per query-group with Bk = 128 and k = 16
What the paper found
MiniMax Sparse Attention, from MiniMax with NVIDIA as a coauthoring affiliation, proposes a blockwise sparse attention layer built on Grouped-Query Attention that lets each GQA group independently select k key-value blocks through a lightweight Index Branch, then applies exact softmax only inside those blocks in the Main Branch. The design is intentionally minimal: a KL alignment loss trains the indexer, while stop-gradient, indexer warmup, and a forced local block stabilize sparse training; at inference, an exp-free TopK kernel and a KV-outer attention kernel turn sparsity into hardware speed. On a 109B-parameter native multimodal MoE model trained on 3T tokens, MSA matches GQA quality while cutting per-token attention compute by 28.4× at 1M context, and with the co-designed kernel it reaches 14.2× prefill and 7.6× decoding wall-clock speedups on H800. The implementation uses Bk = 128 and k = 16, so each query-group attends to at most 2,048 tokens, yet retains strong long-context behavior on RULER and HELMET and stays broadly competitive on MMLU, GSM8K, HumanEval, and multimodal benchmarks such as MMMU, ChartQA, and VideoMME.
Original abstract
Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundreds of thousands to millions of tokens, yet the quadratic cost of softmax attention makes this untenable at deployment scale. We introduce MiniMax Sparse Attention (MSA), a blockwise sparse attention built upon Grouped Query Attention (GQA). A lightweight Index Branch scores key-value blocks and independently selects a Top-k subset for each GQA group, enabling group-specific sparse retrieval while maintaining efficient block-level execution; the Main Branch then performs exact block-sparse attention over only the selected blocks. Designed around a principle of simplicity and scalability, MSA is deliberately streamlined, making it straightforward to deploy efficiently across a broad range of GPUs. To translate sparsity into practical speedups, we co-design MSA with a GPU execution path that uses exp-free Top-k selection and KV-outer sparse attention to improve tensor-core utilization under block-granular access. On a 109B-parameter model with native multimodal training, MSA performs on par with GQA while reducing per-token attention compute by 28.4x at 1M context. Paired with our co-designed kernel, MSA achieves 14.2x prefill and 7.6x decoding wall-clock speedups on H800. Our inference kernel is available at: https://github.com/MiniMax-AI/MSA. A production-grade natively multimodal model powered by MSA has been publicly released at: https://huggingface.co/MiniMaxAI/MiniMax-M3.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.