Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories
AuthorsAli Behrouz, Farnoosh Hashemi, Vahab Mirrokni
Resources
This paper argues that language models should 'sleep' to consolidate memories and generate synthetic practice, helping them keep learning over time without forgetting.
Key results
In few-shot ARC, Sleep reached an 80% success rate, higher than SEAL's 72.5%.
SEAL is reported at 72.5% success rate on the few-shot ARC benchmark.
What the paper found
Language Models Need Sleep Learning to Self Modify and Consolidate Memories proposes a biologically inspired lifecycle for LLMs, built around Google-affiliated authors Ali Behrouz and Vahab Mirrokni, that replaces the usual train/test split with alternating wake and sleep phases. In the wake phase, a Continuum Memory System updates modules at different frequencies; in sleep, the model expands capacity by activating new low-rank experts and performs Knowledge Seeding, an upward self-distillation from a faster, smaller memory block into a larger, slower one using generalized knowledge distillation plus imitation learning with semantic and Levenshtein rewards. A second sleep stage, Dreaming, generates synthetic curriculum data and fine-tunes with ReSTEM-style selection and LoRA to recursively improve the model without external labels. Across Llama-3B/8B and Qwen3 backbones, the method improves class-incremental learning on CLINC, Banking77, and DBpedia, strengthens long-context retrieval on MK-NIAH, LongHealth, and QASPER, and reaches nearly perfect performance on BABILong up to 10M tokens. On factual incorporation from SQuAD, it outperforms SEAL, and on few-shot ARC it reaches an 80% success rate versus 72.5% for SEAL. The key novelty is not just self-distillation, but periodic consolidation into newly expanded parameters followed by controlled synthetic dreaming, which the authors argue reduces catastrophic forgetting while preserving and abstracting newly acquired knowledge.
Original abstract
The past few decades have witnessed significant advances in the design of machine learning algorithms, from early studies on task-specific shallow models to more general deep Large Language Models (LLMs). Despite showing promising results in tasks that require instant prediction or in-context learning, existing models lack the ability to continually learn and effectively transfer their temporal in-context knowledge to their long-term parameters. Inspired by human learning process, we introduce a ''Sleep'' paradigm that allows the models to continually learn, distill their short-term fragile memories into stable long-term knowledge with replay, and recursively improve themselves with ''Dreaming'' process. In more detail, sleep consists of two stages: (1) Memory Consolidation: an upward distillation process, called Knowledge Seeding, where the memories of a smaller-self are distilled into a larger network to provide more capacity while preserving the knowledge. As a proof of concept, we present a new Generalized Distillation process for {Knowledge Seeding} (i.e., the combination of on-policy distillation with Reinforcement Learning (RL)-based imitation learning); (2) Dreaming: a self-improvement phase, where the model uses RL to generate a curriculum of synthetic data to rehearse new knowledge and refine existing capabilities without human supervision. Our experiments on long-horizon, continual learning, knowledge incorporation, and few-shot generalization tasks support the importance of the sleep stage.
Read the original paperMore in Continual Learning
Browse all 24 papers →ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
Haodong Lu, Dong Gong
ASCENT lets deployed LLM agents learn from verified successes on the fly by converting hindsight about their own trajectories into lasting weight updates.
From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang
SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.
Local Support Learning
Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.