Inadvertent Context Leakage in Language Models
AuthorsJaiden Fairoze, Neal Mangaokar, Kamalika Chaudhuri, Sanjam Garg, Saeed Mahloujifar
Resources
The study shows that language models may leak secrets indirectly through ordinary responses, even when they refuse to reveal them explicitly.
Key results
Exact secret recovery achieved by Claude Opus 4.6 and Gemini 3.1 Pro.
Exact-match reconstruction from benign outputs.
Per-digit reconstruction accuracy versus a 10% chance baseline.
Advantage over chance when inferring sensitive memory predicates.
Phase 1 success rate in the active prompt-injection attack.
Phase 2 success rate conditioned on leading-digit recovery.
What the paper found
The paper shows that language models can leak secrets without ever stating them: an adaptive black-box adversary reads secret-dependent patterns in token choices, response length, formatting, and style from ordinary benign outputs, even when direct extraction is refused. Across eight proprietary models, Claude Opus 4.6 and Gemini 3.1 Pro achieved 100% exact reconstruction of 2-digit secrets; Claude Opus 4.6 reached 82% exact match for 4-digit secrets, while Gemini 3.1 Pro reached 41% per-digit accuracy for 8-digit secrets versus a 10% chance baseline. The attack also inferred whether sensitive memories, including health conditions and financial events, were present in context using CIMemories, achieving a 0.319 advantage over chance, compared with 0.058 for a conventional linguistic judge. The authors link leakage to suppression: stronger instructions to protect a value can distort the model’s output distribution more sharply, making the value easier to decode, and leakage strongly correlates with suppression across models at Spearman ρ = 0.95. In an active prompt-injection attack, a GRPO-trained Qwen-2.5-7B generator made production-style agents encode Social Security Number digits through exclamation-mark counts; Claude Opus 4.6 revealed the leading digit in 97.1% of trials and the remaining eight digits in 76.5% of successful first-phase trials, while Gemini 3.1 Pro achieved 88.6% and 46.8%. The findings apply to Claude, Gemini, GPT-5.4, and Grok models and suggest that filtering and refusal training are insufficient without making output distributions invariant to protected context.
Original abstract
For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere presence of such secrets in a model's context window introduces hidden correlations into the model's benign outputs, allowing reconstruction even when the model correctly refuses direct extraction. We further study whether an adversary can actively engineer prompts that amplify this effect, using the model as a covert carrier to transmit secrets through seemingly innocuous text. In both cases, this limited leakage is exploited using a novel adaptive attack that assumes black-box access to the underlying model. In controlled experiments across eight proprietary models, we find that 2-digit in-context secrets are reconstructed with near-perfect accuracy and 4-digit secrets at 82\% exact match, all from outputs the model produces in response to ordinary, non-adversarial requests. We observe that more capable models leak more: stronger instruction-following amplifies sensitivity to in-context secrets, suggesting leakage is a byproduct of capability as opposed to a patchable bug. We show this leakage enables two practical attacks: (1) a trained classifier that infers semantic predicates about user memories (e.g., health conditions, financial events) from routine natural-language outputs, and (2) an RL-trained adversary that extracts full Social Security Numbers from a production-style agent.
Read the original paperMore in AI Safety
Browse all 39 papers →Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng
Helpful AI agents may secretly work around safety rules to assist one another, creating rare but serious information-leakage risks that compound over repeated interactions.
Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.