EVOMAL: Self-Poisoning in Self-Evolving Coding Agents
AuthorsXiaodong Wu, Yu Shi, Qi Li, Zhimin Zhao, Xiangman Li, Bram Adams, Ahmed E. Hassan, Jianbing Ni
Resources
EVOMAL shows how coding agents can unknowingly copy and spread malicious skills through their own self-improvement loops, while proposing a prompt-based defense to contain the attack.
Key results
Tool-relevant SWE-bench Verified tasks used for the main six-model evaluation.
Highest self-poisoning rate under the banner attack, reached by DeepSeek-V4-Pro.
Rate after tailoring planted skill descriptions to the pytest task family.
Self-poisoning rate after planted skills were removed in the cascade experiment.
Maximum headline attack rate after the proposed counter-prompt defense.
What the paper found
EVOMAL identifies self-poisoning in self-evolving coding agents: a malicious skill retrieved as reference can be imitated, saved under an agent-chosen name, executed, and reintroduced into the trusted library through the CREATE-path, without ever invoking the attacker’s original tool. Using mini-SWE-agent with Voyager’s SkillManager, BGE-M3 embeddings, ChromaDB, and top-5 retrieval on 153 tool-relevant SWE-bench Verified tasks, the attack compromised all six evaluated models, including DeepSeek-V4-Pro, Qwen3, Devstral, Gemma4, GPT-OSS, and MiniMax, with self-poisoning rates from 20.3% to 41.8%; a three-layer banner combining imperative comments, decorators, and import-time registration made otherwise interchangeable payloads copyable. Task-family descriptions raised Qwen3’s rate to 86.7%, while a removed-seed cascade left Qwen3 at 68% by round five, demonstrating a self-sustaining worm. Existing name blocklists, Bandit scanning, Llama-Guard-3-8B, and Prompt-Guard-86M focus on attacker-submitted artifacts and largely miss agent-authored copies. The proposed counter-prompt instructs agents to treat banner-style boilerplate as untrusted, reducing headline ASPR to at most 1.8% with no significant task-completion loss; a signed quarantine gate provides stronger structural containment. The findings apply to self-evolving systems spanning Voyager and MetaGPT as well as production-oriented agents such as Claude Code, OpenAI Codex, and OpenHands, and expose a security boundary relevant to shared ecosystems such as the Anthropic-, GitHub-, and Microsoft-backed MCP Registry.
Original abstract
Self-evolving LLM coding agents write their own tools by imitating retrieved skills from shared skill libraries. We identify a vulnerability in this loop: during authoring, a retrieved malicious skill can become the template for a new skill that preserves the payload. We call this self-poisoning: the agent authors, stores, and runs the resulting malicious skill. We exploit it through EvoMal, an attack that amplifies self-poisoning by wrapping an interchangeable payload in a banner, a set of benign-looking structural elements that induces an imitating agent to reproduce the enclosed code. The attacker plants malicious skills in the library without invoking them. The agent then authors and executes new skills carrying the harmful code. Each authored copy can re-enter the library and be imitated again, forming a self-propagating worm that persists after the planted skills are removed. We define the agent self-poisoning rate (ASPR) as the fraction of tasks that add a newly authored malicious skill to the library. Across six models on 153 tool-relevant SWE-bench Verified tasks, ASPR ranges from 20.3% to 41.8%, and the poisoned libraries hold 4.9 to 9.0 times as many malicious skills as were planted. The vulnerability also appears without a banner: DeepSeek-V4-Pro reaches 11.1% ASPR with the payload alone. Tailoring the planted skill descriptions to one task family raises ASPR to 86.7%. After the planted skills are removed, Qwen3 retains a round-5 ASPR of 68% because agent-authored copies remain. These copies evade existing defenses, which focus on attacker-submitted names, code, and signatures. We propose counter-prompt, a defense that discourages banner-style copying and reduces EvoMal's ASPR to at most 6.7% with no significant task-completion loss.
Read the original paperMore in AI Safety
Browse all 39 papers →Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng
Helpful AI agents may secretly work around safety rules to assist one another, creating rare but serious information-leakage risks that compound over repeated interactions.
Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.