Resources
Memory Attention replaces some costly attention computation with reusable token memories, potentially making language models cheaper to run and easier to fit on limited GPU hardware.
Key results
Average zero-shot accuracy improvement for the larger MHA model.
MA retrieval score at 4K tokens, compared with 25.9% for standard attention.
Training-token efficiency at a matched loss.
Training-token efficiency at a matched loss.
Reduction in GPU parameter storage relative to GPU-resident MA.
What the paper found
Memory Attention, or MA, replaces the dedicated value projection in Transformer self-attention with V = K + M, where K is the contextual key representation and M is a layer-specific token embedding retrieved by token ID. RoPE and standard attention weighting remain unchanged, while RMSNorm can be folded into the memory tables at inference, reducing value construction to lookup and addition. The MA-Offload variant stores these tables in CPU memory and prefetches them to the GPU, expanding parameter capacity without keeping all memory parameters on the accelerator. On FineWeb-10BT, matched experiments using 10B and 20B training tokens improved average zero-shot accuracy, with the larger MHA model gaining 1.16 percentage points, although MA also adds parameters and therefore does not isolate architectural effects from capacity gains. On single-needle NIAH retrieval, MA reached a 41.9% mean score at 4K tokens versus 25.9% for standard attention, despite training with a 2K context window. Training efficiency at matched loss improved by 1.42 times for the 1,024-dimensional model and 1.16 times for the 2,048-dimensional model. In a prototype running on an NVIDIA H800, MA-Offload reduced GPU parameter storage by 55.45% relative to GPU-resident MA while retaining comparable forward-pass latency, illustrating a trade-off between lookup and transfer overhead, model capacity, and accelerator memory.
Original abstract
Language models typically construct attention values from contextual hidden states, even when some of their content may be reusable across contexts. We investigate whether token-indexed memory can replace the dedicated value projection when complemented by contextual information. We propose Memory Attention (MA), which forms values by combining layer-specific token memory with contextual keys. The memory supplies token-specific representations, while the keys preserve context dependence. At inference, normalization can be folded into the memory tables, reducing value construction to lookup and addition. Token-indexed retrieval also enables CPU offloading with prefetching, reducing GPU parameter storage. Under matched training token budgets and with additional memory parameters, experiments across attention configurations show improved language modeling and average downstream performance.
Read the original paperMore in Attention Mechanisms
Browse all 18 papers →CoWindow Attention: Full Causal Coverage Is a Collective Property
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang
HLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu
MinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.