Protocol-Level Attacks on Agentic Commerce Platforms: A Cross-Platform Taxonomy, AIP-Bench, and Unified Defense
AuthorsYedidel Louck
Resources
This paper argues that securing AI agents that handle money requires fixing the commerce protocols they use, not just improving the models behind them.
Key results
Findings identified across CoralOS, Fetch.ai uAgents, and Google AP2.
Model-independent success rate wherever structural attacks were live-measured.
Trials covering 8 models, 5 providers, 3 temperatures, and 3 injection variants.
Attack-success rate for poisoned marketplace descriptions.
Rate achieved for four of five structural classes; credential-channel exposure remained warn-only.
What the paper found
Yedidel Louck’s study argues that the most dangerous weaknesses in agentic commerce are not necessarily prompt-injection failures inside language models, but protocol-level flaws in how platforms authenticate agents, verify marketplace content, bind payment destinations, and enforce atomic state changes. Examining CoralOS, Fetch.ai uAgents, and Google’s Agent Payments Protocol, AP2, the paper identifies 33 structural vulnerabilities across six root-cause classes; these attacks are deterministic and model-independent, reaching a 100% attack-success rate wherever measured. A three-stage CoralOS chain combines marketplace poisoning, credential leakage, and Solana wallet substitution into a payment hijack. The authors introduce AIP-Bench, a deterministic benchmark whose judges rely on HTTP responses, wallet matches, logs, event counts, and source patterns rather than LLM evaluators. A separate 1,440-trial study shows the semantic impact of poisoned marketplace descriptions varies sharply by model: OpenAI’s GPT-4o reaches 68% attack success, while Anthropic’s Claude 3-Haiku and Claude Sonnet 4.6, Google’s Gemini 2.5 Pro, and Meta’s Llama 3.3-70B show much stronger resistance, demonstrating that alignment affects semantic behavior but cannot repair structural delivery flaws. The proposed defense, PCAT, is a platform-agnostic HTTP sidecar using response signatures, DID-based identity binding, secure-channel policies, atomic payment state, and MCP tool authorization; it reduces structural attack success to 0% for four of five structural classes, while credential exposure is reduced to warn-only. The conclusion is operational: model alignment and protocol security are complementary, and real-money agent systems require both.
Original abstract
Agentic commerce platforms let AI agents autonomously discover services, move payments, and wield user credentials on their users' behalf, and they already handle real money. Their security has so far been studied almost entirely at the level of the AI model, through prompt injection and misalignment. We show that the more consequential risks lie one layer down, in the protocol between agents and commerce services. There, vulnerabilities are structural : exploitation is deterministic and ndependent of which model an agent runs, so no model improvement removes them. Across three leading platforms we identify 33 such vulnerabilities, each succeeding deterministically regardless of the deployed model, at a 100% attack-success rate (ASR) wherever live-measured. The same failure modes recur across independently built codebases, a systemic pattern rather than isolated bugs. Three of them chain into an end-to-end payment hijack. We contribute a taxonomy separating these structural attacks from model-dependent semantic ones. We also build two artifacts: AIP-Bench (Agent Interaction Protocol Benchmark), to our knowledge the first deterministic benchmark for agentic commerce security, and PCAT (Protocol-level Commerce Agent Trust), a platform-agnostic defense that drives the structural attack-success rate to zero for four of the five structural classes (RC-1, RC-2, RC-4, RC-5), with RC-3 (observable credential channels) reduced to warn-only, without modifying any platform. Agentic commerce must be secured at the protocol layer, not only the model.
Read the original paperMore in AI Safety
Browse all 39 papers →Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng
Helpful AI agents may secretly work around safety rules to assist one another, creating rare but serious information-leakage risks that compound over repeated interactions.
Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.