NTH

Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection

AuthorsYan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao

AffiliationsPeking University · National University of Singapore · BYD Company Limited · Shenzhen University Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model’s own chat template. A forged template marker such as <|im_start|>

October 6, 2026 3 min read
Watch on YouTube
The one-line take

Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.

Key results

58.2
Llama-3.1 identity gap

Percentage-point reduction from reserved-marker authority on InjecAgent direct-harm attacks

49.8
Qwen3-8B suppressed-reasoning gap

Percentage-point identity gap after suppressing the reasoning block

10.3
AgentDojo gap

Percentage-point gap on Qwen3-8B across 409 multi-turn tasks

33
Uncovered tokenizer configurations

Configurations out of 67 that leave reachable tool-protocol tokens intact

255
Affected checkpoints

Popular chat-model checkpoints out of 400 covered by the uncovered configurations

98.4
Llama-3.1 nearest-vector success

Attack success after replacing the reserved marker vector with a nearby ordinary-token vector

What the paper found

This paper shows that chat-template prompt injection depends not only on visible marker text but also on whether the server-side tokenizer emits a reserved control-token ID. The researchers compare byte-identical payloads in which forged markers such as Qwen3’s <|im_start|> remain reserved tokens or are split into ordinary subwords, using token-count-matched controls on InjecAgent and execution-level tests in AgentDojo. On Llama-3.1, GLM-4.5, and Seed-OSS-36B, removing reserved IDs reduces attack success by 39 to 66 percentage points; Llama-3.1 direct-harm success falls from 98.2 percent to 39.7 percent, with a 58.2-percentage-point identity gap. Qwen3-8B is different: its text-based reasoning preserves more of the attack, but suppressing the reasoning block widens the gap to 49.8 percentage points. Vector-swap experiments show that authority is concentrated in the learned embedding at the marker position: on Llama-3.1, a nearby ordinary-token vector restores 98.4 percent success, while averaging subword vectors does not. The effect transfers to multi-turn AgentDojo, producing a 10.3-percentage-point gap on 409 tasks. The standard Hugging Face tokenizer mitigation is incomplete: 33 of 67 configurations covering 255 of 400 popular chat-model checkpoints leave tool-protocol tokens untouched, including configurations associated with DeepSeek and gpt-oss. The findings apply primarily to self-hosted open-weight models such as Qwen3, Llama-3.1, GLM-4.5, and Seed-OSS-36B, rather than hosted OpenAI APIs.

Original abstract

Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as <|im_start|> can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode to exactly the same text, and because tokenization runs on the server, the defender rather than the attacker decides which one the model receives. We use this to measure how much of the injected instruction's authority comes from the reserved token's learned representation. Encoding the forged markers as subwords, with the text held fixed and a control for the extra tokens this adds, lowers attack success on the InjecAgent benchmark by 39 to 66 percentage points on three of four open-weight families, and the gap carries over to multi-turn agent tasks in AgentDojo. On Qwen3-8B the gap is 8 points, because without reserved ids the model still recognises the forged turn from its text by reasoning; suppressing the reasoning block widens the gap to 50. The authority sits in the single learned vector at the marker position: the mean of the marker's subword vectors does not reproduce it, the vector of the nearest ordinary token restores the attack on Llama-3.1, and an adaptive attacker who searches for non-reserved markers finds such embedding neighbours on three of four families. In every base and instruction-tuned pair we test, instruction tuning strengthens the model's preference for reserved markers. The standard mitigation, a tokenizer option that encodes special tokens as ordinary subwords, applies only to tokens a configuration declares special, so in 33 of 67 distinct tokenizer configurations, covering 255 of the 400 most-downloaded chat models on Hugging Face, it leaves intact the tool-protocol tokens through which agents read untrusted tool output, and the gap persists on that channel.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis