NTH

Memory for Large Language Models

AuthorsSining Zhoubian, Dan Zhang, Evgeny Kharlamov, Jie Tang

August 6, 2026 2 min read
Watch on YouTube
The one-line take

This survey explains how modern language models remember, update, and retrieve information, and organizes the field’s many competing memory designs into one coherent framework.

What the paper found

“Memory for Large Language Models,” by Sining Zhoubian, Dan Zhang, Evgeny Kharlamov of Bosch AI, and Jie Tang, proposes a mechanism-centric survey that treats memory as a first-class architectural dimension rather than an incidental result of scaling. Its taxonomy organizes LLM memory along three orthogonal axes: representation, distinguishing implicit computation-coupled states from explicit addressable storage; update dynamics, separating offline training updates from online inference-time adaptation; and persistence, distinguishing short-term context memory from long-term storage. The survey unifies Transformer KV caches, sparse attention, Mamba and other recurrent state-space models, test-time training systems such as Titans, lookup mechanisms such as Engram, and conditional parameter routing in MoE models including Mixtral and DeepSeek-MoE. It highlights a fundamental trade-off: attention provides high-fidelity token recall but has quadratic sequence-cost scaling, while recurrent models offer linear-time processing and bounded storage by compressing history into latent states, at the cost of information loss and interference. Explicit memory improves persistence and controllability through writable parameters, memory slots, or independent lookup stores, but introduces drift, stability, indexing, and lifecycle-management problems. The survey also discusses DeepSeek’s latent KV compression, Apple’s RATTENTION and CommVQ, hybrid systems such as Jamba and Kimi Linear, and evaluation with RULER, LongBench, ∞Bench, and SCROLLS. Its central conclusion is that future LLMs will require adaptive hybrid memory, multi-timescale updates, hardware–algorithm co-design, and evaluation across capacity, recall fidelity, persistence, robustness, and efficiency.

Original abstract

Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an implicit byproduct of computation to a spectrum of explicit, controllable mechanisms. While recent advances introduce diverse strategies---spanning transient attention, recurrent state dynamics, parameter-efficient adaptations, and scalable lookup storage---this rapid evolution has led to a highly fragmented research landscape. In this survey, we present a systematic, architecture-centric taxonomy of memory in LLMs. Our framework characterizes memory along three orthogonal axes: representation (implicit versus explicit), update dynamics (offline versus online), and persistence (short-term versus long-term). We further formalize the granular mechanisms dictating memory writing, routing, state transitions, and consolidation. This unified perspective elucidates the conceptual boundaries between computation-coupled and independently addressable memory, effectively bridging disparate architectural paradigms. Additionally, we critically analyze hybrid memory architectures, system-level efficiency trade-offs, and multi-dimensional evaluation methodologies. By consolidating these scattered advancements into a cohesive framework, this survey charts the trajectory of memory-centric LLM design and provides a principled foundation for future innovations in scalable and adaptive language modeling.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis