Memory for Large Language Models
AuthorsSining Zhoubian, Dan Zhang, Evgeny Kharlamov, Jie Tang
Resources
This survey explains how modern language models remember, update, and retrieve information, and organizes the field’s many competing memory designs into one coherent framework.
What the paper found
“Memory for Large Language Models,” by Sining Zhoubian, Dan Zhang, Evgeny Kharlamov of Bosch AI, and Jie Tang, proposes a mechanism-centric survey that treats memory as a first-class architectural dimension rather than an incidental result of scaling. Its taxonomy organizes LLM memory along three orthogonal axes: representation, distinguishing implicit computation-coupled states from explicit addressable storage; update dynamics, separating offline training updates from online inference-time adaptation; and persistence, distinguishing short-term context memory from long-term storage. The survey unifies Transformer KV caches, sparse attention, Mamba and other recurrent state-space models, test-time training systems such as Titans, lookup mechanisms such as Engram, and conditional parameter routing in MoE models including Mixtral and DeepSeek-MoE. It highlights a fundamental trade-off: attention provides high-fidelity token recall but has quadratic sequence-cost scaling, while recurrent models offer linear-time processing and bounded storage by compressing history into latent states, at the cost of information loss and interference. Explicit memory improves persistence and controllability through writable parameters, memory slots, or independent lookup stores, but introduces drift, stability, indexing, and lifecycle-management problems. The survey also discusses DeepSeek’s latent KV compression, Apple’s RATTENTION and CommVQ, hybrid systems such as Jamba and Kimi Linear, and evaluation with RULER, LongBench, ∞Bench, and SCROLLS. Its central conclusion is that future LLMs will require adaptive hybrid memory, multi-timescale updates, hardware–algorithm co-design, and evaluation across capacity, recall fidelity, persistence, robustness, and efficiency.
Original abstract
Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an implicit byproduct of computation to a spectrum of explicit, controllable mechanisms. While recent advances introduce diverse strategies---spanning transient attention, recurrent state dynamics, parameter-efficient adaptations, and scalable lookup storage---this rapid evolution has led to a highly fragmented research landscape. In this survey, we present a systematic, architecture-centric taxonomy of memory in LLMs. Our framework characterizes memory along three orthogonal axes: representation (implicit versus explicit), update dynamics (offline versus online), and persistence (short-term versus long-term). We further formalize the granular mechanisms dictating memory writing, routing, state transitions, and consolidation. This unified perspective elucidates the conceptual boundaries between computation-coupled and independently addressable memory, effectively bridging disparate architectural paradigms. Additionally, we critically analyze hybrid memory architectures, system-level efficiency trade-offs, and multi-dimensional evaluation methodologies. By consolidating these scattered advancements into a cohesive framework, this survey charts the trajectory of memory-centric LLM design and provides a principled foundation for future innovations in scalable and adaptive language modeling.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.