Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation
AuthorsEmilio Ferrara
Resources
A large study finds that open-weight language models can contain information about changes to their own internals but generally cannot reliably tell us about those changes.
Key results
Total measurements across eight open-weight models
Intervention-versus-sham discrimination, effectively chance
Upper bound in percentage points on practical AUROC advantage
Highest held-out intervention-presence accuracy from activations
Approximate held-out AUROC after fine-tuning Qwen2.5-7B-Instruct to report interventions
Qwen2.5-7B-Instruct confidence discrimination despite binary reports at chance
What the paper found
The paper introduces Open-Weight Masked Introspection, or OWMI, a framework that intervenes in residual-stream sites, attention heads, and sparse-autoencoder features, then tests whether a language model can report the alteration. Across 78000 measurements on eight open-weight models, including Qwen, Mistral, Llama, Gemma, DeepSeek, Phi, and GLM systems ranging from 0.5B to 15B parameters, reported detection was indistinguishable from guessing: pooled AUROC was approximately 0.5007, with an equivalence test bounding any practical advantage below 0.15 percentage points of AUROC. OWMI controls for sham runs, impact-matched random perturbations, and a text-only observer, while preserving tasks from benchmarks such as MMLU, GSM8K, HumanEval, and TruthfulQA. The failure is not caused by missing information: linear probes recover intervention presence from the same activations at 95.8% and 75.0% held-out accuracy, and the signal remains decodable with no held-out error at tested downstream layers. A LoRA-fine-tuned Qwen2.5-7B-Instruct model validates the instrument, reaching AUROC approximately 1.0 on held-out interventions. One exception appears in confidence rather than words: Qwen2.5-7B-Instruct’s binary report stays at AUROC 0.500, but its confidence separates intervention from sham at AUROC 0.647. The conclusion is that current tested open-weight models contain intervention information internally but do not reliably route it into verbal testimony, challenging chain-of-thought monitoring, self-critique, and confidence elicitation as oversight methods unless they are checked against internal activation references.
Original abstract
Are frontier models able to introspect about their internal states? Recent work suggests that under certain conditions a complex enough model can audit its own internals, call out what changed, and report back confidently about it. We tested that claim on eight open-weight models from seven families and found no such ability: asked whether their own computation had been altered, none answered better than chance. To test it we built Open-Weight Masked Introspection (OWMI), a framework that intervenes on residual-stream sites, attention heads and sparse-autoencoder features, then interrogates the model about the change against the null conditions an answer has to beat: sham runs where nothing was altered, impact-matched random perturbations, and a text-only observer that sees only the visible output. Over 78,000 measurements, no model's report discriminates a real intervention from a sham beyond chance (AUROC ~0.5007), and an equivalence test bounds the effect below 0.15 percentage points of AUROC. Surprisingly, all the information needed is in the models. A model fine-tuned to report this class of intervention reaches near-perfect recovery on held-out directions, and a linear probe recovers intervention presence from the same activations at 75% to 95.8% accuracy, sharpening to no held-out error at the last layer before the model speaks. In one model the signal surfaces in the confidence rather than the words: its yes-or-no report never varies, while the confidence attached to it separates intervention from sham at AUROC 0.647. The failure sits in the path from internal state to verbal report, so oversight that reads a model's own testimony needs validating against an internal reference. While our results show the inability of current open-weight models to introspect, the debate is not settled for future models.
Read the original paperMore in AI Safety
Browse all 39 papers →Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng
Helpful AI agents may secretly work around safety rules to assist one another, creating rare but serious information-leakage risks that compound over repeated interactions.
Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.