NTH

EVOMAL: Self-Poisoning in Self-Evolving Coding Agents

AuthorsXiaodong Wu, Yu Shi, Qi Li, Zhimin Zhao, Xiangman Li, Bram Adams, Ahmed E. Hassan, Jianbing Ni

August 30, 2026 3 min read
Watch on YouTube
The one-line take

EVOMAL shows how coding agents can unknowingly copy and spread malicious skills through their own self-improvement loops, while proposing a prompt-based defense to contain the attack.

Key results

153
Evaluation task count

Tool-relevant SWE-bench Verified tasks used for the main six-model evaluation.

41.8%
Maximum generic ASPR

Highest self-poisoning rate under the banner attack, reached by DeepSeek-V4-Pro.

86.7%
Targeted Qwen3 ASPR

Rate after tailoring planted skill descriptions to the pytest task family.

68%
Qwen3 round-five ASPR

Self-poisoning rate after planted skills were removed in the cascade experiment.

1.8%
Counter-prompt ASPR

Maximum headline attack rate after the proposed counter-prompt defense.

What the paper found

EVOMAL identifies self-poisoning in self-evolving coding agents: a malicious skill retrieved as reference can be imitated, saved under an agent-chosen name, executed, and reintroduced into the trusted library through the CREATE-path, without ever invoking the attacker’s original tool. Using mini-SWE-agent with Voyager’s SkillManager, BGE-M3 embeddings, ChromaDB, and top-5 retrieval on 153 tool-relevant SWE-bench Verified tasks, the attack compromised all six evaluated models, including DeepSeek-V4-Pro, Qwen3, Devstral, Gemma4, GPT-OSS, and MiniMax, with self-poisoning rates from 20.3% to 41.8%; a three-layer banner combining imperative comments, decorators, and import-time registration made otherwise interchangeable payloads copyable. Task-family descriptions raised Qwen3’s rate to 86.7%, while a removed-seed cascade left Qwen3 at 68% by round five, demonstrating a self-sustaining worm. Existing name blocklists, Bandit scanning, Llama-Guard-3-8B, and Prompt-Guard-86M focus on attacker-submitted artifacts and largely miss agent-authored copies. The proposed counter-prompt instructs agents to treat banner-style boilerplate as untrusted, reducing headline ASPR to at most 1.8% with no significant task-completion loss; a signed quarantine gate provides stronger structural containment. The findings apply to self-evolving systems spanning Voyager and MetaGPT as well as production-oriented agents such as Claude Code, OpenAI Codex, and OpenHands, and expose a security boundary relevant to shared ecosystems such as the Anthropic-, GitHub-, and Microsoft-backed MCP Registry.

Original abstract

Self-evolving LLM coding agents write their own tools by imitating retrieved skills from shared skill libraries. We identify a vulnerability in this loop: during authoring, a retrieved malicious skill can become the template for a new skill that preserves the payload. We call this self-poisoning: the agent authors, stores, and runs the resulting malicious skill. We exploit it through EvoMal, an attack that amplifies self-poisoning by wrapping an interchangeable payload in a banner, a set of benign-looking structural elements that induces an imitating agent to reproduce the enclosed code. The attacker plants malicious skills in the library without invoking them. The agent then authors and executes new skills carrying the harmful code. Each authored copy can re-enter the library and be imitated again, forming a self-propagating worm that persists after the planted skills are removed. We define the agent self-poisoning rate (ASPR) as the fraction of tasks that add a newly authored malicious skill to the library. Across six models on 153 tool-relevant SWE-bench Verified tasks, ASPR ranges from 20.3% to 41.8%, and the poisoned libraries hold 4.9 to 9.0 times as many malicious skills as were planted. The vulnerability also appears without a banner: DeepSeek-V4-Pro reaches 11.1% ASPR with the payload alone. Tailoring the planted skill descriptions to one task family raises ASPR to 86.7%. After the planted skills are removed, Qwen3 retains a round-5 ASPR of 68% because agent-authored copies remain. These copies evade existing defenses, which focus on attacker-submitted names, code, and signatures. We propose counter-prompt, a defense that discourages banner-style copying and reduces EvoMal's ASPR to at most 6.7% with no significant task-completion loss.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis