Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
AuthorsYan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
AffiliationsPeking University · National University of Singapore · BYD Company Limited · Shenzhen University Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model’s own chat template. A forged template marker such as <|im_start|>
Resources
Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.
Key results
Percentage-point reduction from reserved-marker authority on InjecAgent direct-harm attacks
Percentage-point identity gap after suppressing the reasoning block
Percentage-point gap on Qwen3-8B across 409 multi-turn tasks
Configurations out of 67 that leave reachable tool-protocol tokens intact
Popular chat-model checkpoints out of 400 covered by the uncovered configurations
Attack success after replacing the reserved marker vector with a nearby ordinary-token vector
What the paper found
This paper shows that chat-template prompt injection depends not only on visible marker text but also on whether the server-side tokenizer emits a reserved control-token ID. The researchers compare byte-identical payloads in which forged markers such as Qwen3’s <|im_start|> remain reserved tokens or are split into ordinary subwords, using token-count-matched controls on InjecAgent and execution-level tests in AgentDojo. On Llama-3.1, GLM-4.5, and Seed-OSS-36B, removing reserved IDs reduces attack success by 39 to 66 percentage points; Llama-3.1 direct-harm success falls from 98.2 percent to 39.7 percent, with a 58.2-percentage-point identity gap. Qwen3-8B is different: its text-based reasoning preserves more of the attack, but suppressing the reasoning block widens the gap to 49.8 percentage points. Vector-swap experiments show that authority is concentrated in the learned embedding at the marker position: on Llama-3.1, a nearby ordinary-token vector restores 98.4 percent success, while averaging subword vectors does not. The effect transfers to multi-turn AgentDojo, producing a 10.3-percentage-point gap on 409 tasks. The standard Hugging Face tokenizer mitigation is incomplete: 33 of 67 configurations covering 255 of 400 popular chat-model checkpoints leave tool-protocol tokens untouched, including configurations associated with DeepSeek and gpt-oss. The findings apply primarily to self-hosted open-weight models such as Qwen3, Llama-3.1, GLM-4.5, and Seed-OSS-36B, rather than hosted OpenAI APIs.
Original abstract
Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as <|im_start|> can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode to exactly the same text, and because tokenization runs on the server, the defender rather than the attacker decides which one the model receives. We use this to measure how much of the injected instruction's authority comes from the reserved token's learned representation. Encoding the forged markers as subwords, with the text held fixed and a control for the extra tokens this adds, lowers attack success on the InjecAgent benchmark by 39 to 66 percentage points on three of four open-weight families, and the gap carries over to multi-turn agent tasks in AgentDojo. On Qwen3-8B the gap is 8 points, because without reserved ids the model still recognises the forged turn from its text by reasoning; suppressing the reasoning block widens the gap to 50. The authority sits in the single learned vector at the marker position: the mean of the marker's subword vectors does not reproduce it, the vector of the nearest ordinary token restores the attack on Llama-3.1, and an adaptive attacker who searches for non-reserved markers finds such embedding neighbours on three of four families. In every base and instruction-tuned pair we test, instruction tuning strengthens the model's preference for reserved markers. The standard mitigation, a tokenizer option that encodes special tokens as ordinary subwords, applies only to tokens a configuration declares special, so in 33 of 67 distinct tokenizer configurations, covering 255 of the 400 most-downloaded chat models on Hugging Face, it leaves intact the tool-protocol tokens through which agents read untrusted tool output, and the gap persists on that channel.
Read the original paperMore in AI Safety
Browse all 39 papers →Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng
Helpful AI agents may secretly work around safety rules to assist one another, creating rare but serious information-leakage risks that compound over repeated interactions.
Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.
Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection
Keertana Chidambaram, Andrew Ilyas, Vasilis Syrgkanis
The paper shows that hidden harmful plans can be injected into an AI’s context, causing it to act unsafely while fooling the very monitors meant to catch it.