NTH

AI Agents Enable Adaptive Computer Worms

AuthorsJonas Guan, Tom Blanchard, Hanna Foerster, Hengrui Jia, Gabriel Huang, Nicolas Papernot

June 4, 2026 3 min read
Watch on YouTube
The one-line take

This paper argues that AI agents can power a new kind of self-sustaining worm that adapts to each target, spreads across real networks, and uses stolen compute to keep attacking at near-zero marginal cost.

Key results

15
autonomous runs

The worm was evaluated in 15 fully autonomous experiments on an isolated heterogeneous network.

33
network size

The testbed contained 33 virtual machines spanning Linux, Windows, and IoT systems.

82%
vulnerability detection

Across all attempts, the agent correctly identified the target vulnerability in 82% of cases.

44%
privileged exploitation

Across all attempts, the agent achieved root/admin/SYSTEM access in 44% of exploitation attempts.

88%
self-replication

On successfully exploited targets, the worm replicated itself in 88% of cases.

20.4
average compromised hosts

Over 7 days of autonomous operation per run, the worm compromised 20.4 hosts on average.

What the paper found

This paper, from the University of Toronto, the Vector Institute, the University of Cambridge, and ServiceNow, demonstrates a proof-of-concept AI-driven computer worm that uses a locally hosted open-weight LLM on a single GPU to generate target-specific attack logic at runtime, rather than relying on fixed exploit code. The key technical contribution is an agentic harness built around phase-structured execution, hierarchical memory, a reasoning graph, tool wrappers, and multi-agent swarm coordination, which compensates for the brittleness of smaller models by feeding them just-in-time exploit guidance and preserving context across steps. In 15 fully autonomous runs on an isolated 33-host heterogeneous network spanning Linux, Windows, and IoT systems, the worm detected vulnerabilities in 82 percent of attempts, achieved privileged exploitation in 44 percent, and successfully self-replicated in 88 percent of exploited cases; overall it compromised 20.4 hosts on average and reached up to 7 generations of propagation over 7 days. The authors show that the system can operationalize newly disclosed 2026 vulnerabilities, including Marimo CVE-2026-39987, Copy Fail CVE-2026-31431, and Dirty Frag CVE-2026-43284/CVE-2026-43500, using runtime advisory data despite the model’s training cutoff. Their central finding is economic as much as technical: because each compromised machine can supply compute for inference, the attacker’s marginal cost per new infection approaches zero, and because the worm runs on open weights rather than a commercial API, centralized safety controls like refusals or rate limits do not meaningfully constrain it.

Original abstract

A computer worm is malware that spreads on a network by replicating itself from one machine to another. Traditional worms, like WannaCry, exploited predetermined vulnerabilities, and their spread can be halted by patching those vulnerabilities. Here we show that artificial intelligence (AI) agents enable a fundamentally new threat: a worm that generates tailored attack strategies to each target it encounters. The worm parasitically uses compromised machines to run open-weight large language models (LLMs) to sustain its reasoning, or extend its reach for further attacks. Deployed on a network of machines spanning Linux, Windows, and IoT (Internet of Things) devices, the worm propagated by exploiting common, real-world corporate network vulnerabilities. Since the worm is powered by stolen compute, the attacker's marginal cost per new infection is zero. This creates a destabilizing economic asymmetry between attackers and defenders. Moreover, because the worm requires no commercial AI platform, centralized safety controls, such as service refusals or rate limiting, are structurally irrelevant. Our results demonstrate that self-sustaining AI-driven cyber-threats are no longer theoretical. We must prepare for autonomous generative adversaries: malware systems that propagate without human operators and are defined not by fixed exploit code, but by the capacity to reason about targets, adapt to observations, and synthesize attack logic in real time.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis