NTH

Memory Attention

AuthorsJiale Kang

September 28, 2026 2 min read
Watch on YouTube
The one-line take

Memory Attention replaces some costly attention computation with reusable token memories, potentially making language models cheaper to run and easier to fit on limited GPU hardware.

Key results

1.16%
Large-model average accuracy gain

Average zero-shot accuracy improvement for the larger MHA model.

41.9%
NIAH 4K mean score

MA retrieval score at 4K tokens, compared with 25.9% for standard attention.

1.42
L24-D1024 token efficiency

Training-token efficiency at a matched loss.

1.16
L24-D2048 token efficiency

Training-token efficiency at a matched loss.

55.45%
MA-Offload GPU storage reduction

Reduction in GPU parameter storage relative to GPU-resident MA.

What the paper found

Memory Attention, or MA, replaces the dedicated value projection in Transformer self-attention with V = K + M, where K is the contextual key representation and M is a layer-specific token embedding retrieved by token ID. RoPE and standard attention weighting remain unchanged, while RMSNorm can be folded into the memory tables at inference, reducing value construction to lookup and addition. The MA-Offload variant stores these tables in CPU memory and prefetches them to the GPU, expanding parameter capacity without keeping all memory parameters on the accelerator. On FineWeb-10BT, matched experiments using 10B and 20B training tokens improved average zero-shot accuracy, with the larger MHA model gaining 1.16 percentage points, although MA also adds parameters and therefore does not isolate architectural effects from capacity gains. On single-needle NIAH retrieval, MA reached a 41.9% mean score at 4K tokens versus 25.9% for standard attention, despite training with a 2K context window. Training efficiency at matched loss improved by 1.42 times for the 1,024-dimensional model and 1.16 times for the 2,048-dimensional model. In a prototype running on an NVIDIA H800, MA-Offload reduced GPU parameter storage by 55.45% relative to GPU-resident MA while retaining comparable forward-pass latency, illustrating a trade-off between lookup and transfer overhead, model capacity, and accelerator memory.

Original abstract

Language models typically construct attention values from contextual hidden states, even when some of their content may be reusable across contexts. We investigate whether token-indexed memory can replace the dedicated value projection when complemented by contextual information. We propose Memory Attention (MA), which forms values by combining layer-specific token memory with contextual keys. The memory supplies token-specific representations, while the keys preserve context dependence. At inference, normalization can be folded into the memory tables, reducing value construction to lookup and addition. Token-indexed retrieval also enables CPU offloading with prefetching, reducing GPU parameter storage. Under matched training token budgets and with additional memory parameters, experiments across attention configurations show improved language modeling and average downstream performance.

Read the original paper

More in Attention Mechanisms

Browse all 18 papers →
01Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis