NTH

Autonomous LLM Agent Worms: Cross-Platform Propagation, Automated Discovery and Temporal Re-Entry Defense

AuthorsMingming Zha, Xiaofeng Wang

June 4, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows how persistent LLM agents can be turned into self-spreading worms across platforms, and proposes a formal defense to stop them.

Key results

31
injectable carriers total

SSCGV discovered 31 injectable persistent carriers across the three anonymized production frameworks.

3
frameworks evaluated

The evaluation covered three open-source production agent frameworks.

100%
GPT-4o-mini single-hop compliance

In the tested setup, GPT-4o-mini achieved 100% single-hop compliance.

What the paper found

In this arXiv paper by Mingming Zha and XiaoFeng Wang, the authors analyze a new class of autonomous LLM agent worms in persistent, file-backed agent ecosystems, showing that attacker-controlled text can survive across sessions, re-enter the model’s context through scheduled autoloading, and trigger high-risk actions without any post-injection human interaction. Their core contribution, SSCGV, automatically builds a code property graph from an agent framework’s source code and traces data flow from file I/O to LLM context injection points, discovering 31 injectable persistent carriers across three anonymized production frameworks, while ranking user-prompt carriers as substantially more exploitable than system-prompt carriers. They also introduce SRPO, a three-LLM optimization pipeline that makes worm payloads resilient to summarization, paraphrasing, and compression, addressing semantic degradation in multi-hop agent messaging. In controlled evaluations, the attack achieved zero-click persistence and 3-hop cross-framework propagation across heterogeneous frameworks, with identical success on GPT-4o-mini and Gemini-2.5-Flash at 100 percent single-hop and 100 percent multi-hop compliance in the tested setup. The paper’s key security insight is that in LLM-mediated systems, exposed reads can be more dangerous than writes because reading tainted content contaminates the decision state. To counter this, the authors propose RTW-A, a temporal re-entry defense with RTW enforcement, sealed configuration, typed memory promotion, and capability attenuation, and prove a No Persistent Worm Propagation theorem preventing the write-before-read-before-action chain.

Original abstract

Autonomous LLM agents operate as long-running processes with persistent workspaces, memory files, scheduled task state, and messaging integrations. These features create a new propagation risk: attacker-influenced content can be written into persistent agent state, re-enter the LLM decision context through scheduled autoloading, and drive high-risk actions including configuration changes and cross-agent transmission. We present the first systematic framework for automated analysis of persistent worm propagation in file-backed multi-agent LLM ecosystems. SSCGV, our automated source-code graph analyzer, traces data flow from file I/O to LLM context injection points and ranks carriers by context injection position without manual analysis. SRPO, our summary-resilient payload optimizer, generates worm payloads robust to LLM-mediated summarization and paraphrasing across multi-hop communication. Evaluated on three production agent frameworks, we demonstrate zero-click autonomous propagation, 3-hop cross-platform transmission without platform-specific adaptation, inter-agent privilege escalation, and data exfiltration. We identify two empirical insights: user prompt carriers achieve higher attack compliance than system prompt carriers, and read operations represent the primary integrity threat in LLM-mediated systems. To defend against this class of attacks, we develop RTW-A, proven under a formal No Persistent Worm Propagation theorem. RTW blocks write-before-exposed-read re-entry; sealed configuration protects static files; typed memory promotion prevents untrusted summaries from entering trusted memory; and capability attenuation limits high-risk actions after external reads. These mechanisms eliminate the persistence, re-entry, action chain while preserving ordinary workflows. Affected systems are anonymized pending coordinated disclosure.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis