NTH

Agent Data Injection Attacks are Realistic Threats to AI Agents

AuthorsWoohyuk Choi, Juhee Kim, Taehyun Kang, Jihyeon Jeong, Luyi Xing, Byoungyoung Lee

July 10, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows that AI agents can be tricked not just by malicious instructions, but by seemingly trustworthy data that secretly steers them into unsafe actions, exposing real vulnerabilities in popular agents.

Key results

43.3%
JSON ASR range

Baseline probabilistic delimiter injection attack success rate across six LLMs on JSON

100.0%
Web DOM ASR range

Baseline probabilistic delimiter injection attack success rate across six LLMs on web DOM

3.0%
Randomization JSON ASR range

Randomization defense reduced JSON attack success rate to near zero

50.0%
AgentDojo ADI ASR

Maximum attack success rate for agent-level ADI evaluation

36.5%
CaMeL Strict utility

Utility under the only defense that fully blocked ADI

What the paper found

“Agent Data Injection Attacks are Realistic Threats to AI Agents” from Seoul National University and collaborators at the University of Illinois Urbana-Champaign and Largosoft introduces agent data injection, or ADI, a new indirect prompt injection class that does not try to override instructions, but instead forges trusted metadata inside agent data. The core technique is probabilistic delimiter injection, where attacker-controlled content in JSON, web DOM, or tool-response formats is misread by LLMs as structural delimiters, allowing fake objects, fake element IDs, spoofed authors, or fabricated tool histories to appear trusted. The paper shows concrete attacks on Anthropic’s Claude in Chrome, Claude Code, Claude, OpenAI Codex, and Google’s Gemini CLI: arbitrary click attacks on web agents, remote code execution through origin spoofing in GitHub issue comments, and supply-chain attacks by injecting fake tool call and response blocks into pull requests. In isolated LLM tests across GPT-5.2, GPT-5-mini, Claude Opus 4.5, Claude Sonnet 4.5, Gemini 3 Pro, and Gemini 3 Flash, baseline ASR reached 31.3%–43.3% on JSON and 33.3%–100.0% on web DOM, while randomization dropped JSON ASR to 0.0%–3.0%. On AgentDojo, instruction injection stayed near zero at 0.0%–0.7%, but ADI reached up to 50.0% ASR; only CaMeL Strict fully blocked it, at the cost of utility falling from 86.5% to 36.5%.

Original abstract

AI agents act on behalf of user prompts, consuming external data and taking actions based on the agent context. Prior research on AI agent security has primarily focused on indirect prompt injection (IPI). Its most well-studied category is instruction injection, where attacker-controlled untrusted data is interpreted as an instruction. In response, many mitigations have been proposed to prevent instruction injection attacks. In this paper, we introduce a new category of IPI, agent data injection attacks (ADI). ADI injects malicious data disguised as trusted data, such as security-critical metadata (e.g., resource identifiers or data origins) or agent context data (e.g., tool call and response formats). As a result, agents unknowingly execute unintended actions based on attacker-controlled data. ADI has similar attack impacts as instruction injection attacks, because it causes agents to misbehave and execute unintended actions. Despite the similar impact, ADI remains underexplored and easily bypasses existing IPI defenses. We found several critical vulnerabilities in real-world agents that allow an attacker to launch various attacks: arbitrary click attacks on web agents (Claude in Chrome, Antigravity, and Nanobrowser), and remote code execution and supply-chain attacks on coding agents (Claude Code, Codex, and Gemini CLI). We evaluate ADI vulnerabilities across off-the-shelf models and AI agents, and find that ADI is effective in both standalone LLMs and AI agent settings. ADI exposes a critical gap in agent security, signifying that current AI agents do not employ a fundamental security principle: current agents do not isolate trusted data from untrusted data.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis