Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability
AuthorsAnkit Goyal, Jaideep Ray
Resources
When an AI agent gets a new model, structured memories survive far better than compressed notes or partially migrated retrieval indexes.
Key results
Synthetic histories used for controlled migration experiments
Accuracy change after replacing the memory writer, reported as +0.0004 ± 0.0020
Accuracy-point gain from upgrading bge-large-en from v1.0 to v1.5 with full re-embedding
Accuracy-point gain from a 50/50 old/new embedding index
Share of NOTES accuracy deficit attributed to initial compression
Histories recovered to the 90% target out of 48 using Qwen and retained raw history
What the paper found
This controlled study asks whether an AI agent’s long-term memory survives a model upgrade rather than merely a database copy. Across 48 synthetic histories and 160 exact-scored questions per history, it compares full-context raw history, retrieval-augmented generation, model-compressed NOTES, and a fixed-schema knowledge graph using Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct-1M. The key result is that structure is more portable than prose: KG-fixed accuracy changed by only +0.0004 ± 0.0020 after a writer swap, while NOTES shifted asymmetrically by +9.91 or −13.28 percentage points depending on migration direction. Retrieval has a separate failure mode: upgrading BAAI/bge-large-en from v1.0 to v1.5 produced an 11.90-point gain after full re-embedding, but a 50/50 mixed index gained only 4.96 points because vectors from different embedding spaces silently damage ranking, even when dimensions match. Diagnostics show that 80% of NOTES loss occurs during initial compression, whereas 81% of RAG loss comes from retrieval misses before the reader sees the evidence. Store-only NOTES repair reached the 90% recovery target in 0 of 48 cases; retaining raw histories enabled Qwen to recover 34 of 48 cases, at a median cost of $0.76. The findings are relevant to production memory systems, including Claude-style import and export workflows: test each writer-to-reader direction, isolate embedding versions, preserve provenance and source histories when policy permits, and diagnose writing, retrieval, and reading separately.
Original abstract
Model upgrades are routine; memory migrations are not. An agent can keep the same memory store and still forget: a new model may interpret old notes differently, mixed embedding versions may break retrieval, and repair may fail without the original evidence. We compare memory as the same history is preserved verbatim for long-context reading (LC-RAW), divided into chunks for retrieval-augmented generation (RAG), compressed by a model into natural-language notes (NOTES), or normalized into a fixed-schema knowledge graph (KG-fixed). The study uses 48 synthetic histories with randomized answer codes, exact scoring, and two open-weight models with sub 10 billion parameters. Our measurements show that fixed-schema structures transfer reliably, with KG-fixed accuracy changing by only $+0.0004 \pm 0.0020$ following a writer swap. Conversely, compressed NOTES exhibit high model coupling, with accuracy shifting asymmetrically by $+9.91$ or $-13.28$ percentage points depending on the specific migration direction. In RAG systems, partial embedding migrations using a 50/50 mixed index capture only a 4.96-point accuracy improvement, forfeiting the majority of the 11.90-point gain achieved through full re-embedding. Diagnostic decomposition attributes 80% ($0.467 \pm 0.014$) of the NOTES accuracy deficit to information lost during initial construction, whereas retrieval failures drive 81% ($0.364 \pm 0.012$) of the RAG deficit. Finally, store-only repair of NOTES fails to reach a 90% performance recovery target in all 48 test cases, whereas retaining the raw source history enables successful recovery in 34 of 48 cases for one tested direction. These findings highlight the necessity of direction-specific migration testing, strict embedding space isolation, and the retention of source histories for memory repair.
Read the original paperMore in Continual Learning
Browse all 24 papers →ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
Haodong Lu, Dong Gong
ASCENT lets deployed LLM agents learn from verified successes on the fly by converting hindsight about their own trajectories into lasting weight updates.
From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang
SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.
Local Support Learning
Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.