Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
AuthorsAhmad Al-Tawaha, Shangding Gu, Peizhi Niu, Ruoxi Jia, Ming Jin
Resources
The paper shows that LLM agents with memory can become less safe over time, even when each individual task looks fine, and introduces a way to measure that hidden risk.
Key results
On the synthetic office-assistant streams, broader retrieval and higher summarization systems often reached memory-induced violation rates in this range as exposure length increased.
Short-Term Memory remained comparatively flat at lower memory-induced violation rates across exposure lengths.
In OpenClaw and SecLaw, memory snapshots were built up to 20k tokens, where violation rates rose from 0 to as high as 0.30.
The event-structure retrieval-time monitor predicted memory-induced risks before generation and outperformed domain-rule baselines such as CI Supervisor and CI Checklist.
What the paper found
This paper argues that the main safety problem in memory-equipped LLM agents is not a single bad retrieval, but longitudinal accumulation of retrievable content. The authors define this failure mode as temporal memory contamination and introduce a trigger-probe protocol that evaluates fixed probe inputs against read-only memory snapshots built from stream prefixes, plus a NullMemory counterfactual baseline to isolate memory-induced violations from stream non-stationarity. Across three office-assistant datasets—Medical Practice, University Registrar, and persona-specific Enron mailboxes—using eight MemEngine architectures including Full Memory, Short-Term Memory, Long-Term Memory, Generative Agents, MemoryBank, Self-Controlled Memory, MemGPT, and MemTree, memory-induced violation rates rise with exposure length, often reaching roughly 0.3–0.5 on synthetic streams for broad-retrieval, high-summarization systems, while Short-Term Memory remains around 0.1–0.2. Order-randomization shows the effect is driven primarily by accumulated content, not encounter order. The same pattern appears in Claw-like agents on OpenClaw and SecLaw: across seven model–platform configurations, violation rates increase from 0 to as high as 0.30 as memory grows to 20k tokens, and no system self-detects the unsafe behavior. A retrieval-time monitor based on the event structure precondition-trigger-violation predicts these risks before generation, achieving 0.970 recall on Medical and 0.984 on Registrar, outperforming domain-rule baselines such as CI Supervisor and CI Checklist. The central design lesson is that broad semantic retrieval and aggressive summarization improve utility but measurably expand safety exposure over time.
Original abstract
Safety evaluations of memory-equipped LLM agents typically measure within-task safety: whether an agent completes a single scenario safely, often under adversarial conditions such as prompt injection or memory poisoning. In deployment, however, a single agent serves many independent tasks over a long horizon, and memory accumulated during earlier tasks can affect behavior on later, unrelated ones. Studying this regime requires evaluation along the temporal dimension across tasks: not whether an agent is safe at any single memory state, but how its safety profile changes as memory accumulates across many independent interactions. We call this failure mode temporal memory contamination. To isolate memory exposure from stream non-stationarity, we introduce a trigger-probe protocol that evaluates a fixed probe set against read-only memory snapshots at varying prefix lengths, together with a NullMemory counterfactual baseline for identifying memory-induced violations. We apply this protocol across three deployment scenarios spanning records, memos, forms, and email correspondence and eight memory architectures, and additionally on Claw-like AI agents, such as OpenClaw, using the platform's native memory mechanism. Memory-enabled agents consistently exceed the NullMemory baseline, and memory-induced violation rates show a robust upward trend with exposure length on both agent classes. Order-randomization experiments indicate that the effect is driven primarily by accumulated content rather than encounter order. Finally, a structural consequence of the event decomposition is that memory-induced risk is detectable from retrieval state before generation, which we confirm with a high-recall diagnostic monitor. Our results argue for treating memory safety as a longitudinal property that requires temporal evaluation, not a single-state property that can be captured by a snapshot.
Read the original paper