EchoCoT: Extracting Hidden Chain-of-Thought from Large Reasoning Models
AuthorsYiting Qu, Ziqing Yang, Chi Cui, Ye Leng, Junjie Chu, Yang Zhang
Resources
EchoCoT shows that hidden reasoning traces from powerful AI models may be recoverable through carefully designed API interactions, creating a major security concern.
Key results
Maximum near-verbatim extraction success achieved by EchoCoT-LTGO across DeepSeek-V4-Flash, Qwen3.5-Plus, and GLM-5.2.
Maximum transfer success on MATH500, JEEBench, and LiveCodeBench.
Tokens extracted by EchoCoT from Gemini-2.5.
Provider-reported target reasoning-token count for the Gemini-2.5 example.
Mean attack success rate after applying the tested defensive system prompt.
What the paper found
EchoCoT demonstrates that hidden chain-of-thought can be recovered from black-box reasoning models through a tool-call replay surface: while ordinary multi-turn interactions discard internal reasoning, tool calls preserve it within the same turn. The attack repeatedly asks a scratchpad tool to archive the model’s reasoning, evaluates each candidate using API-reported reasoning-token counts and compressed CoT summaries, and adapts subsequent injections. Its LLM-based INJECT-REFLECT-DISTILL optimizer searches universal trajectories using length-guided optimization, or LTGO when textual summaries are available. Across DeepSeek-V4-Flash, Qwen3.5-Plus, and GLM-5.2, EchoCoT-LTGO reaches up to 66.4% ASR@90, requiring both length error no greater than 10% and Token-EM of at least 0.90, and transfers to MATH500, JEEBench, and LiveCodeBench with success rates up to 80%. The method also produces long traces: it extracts 33463 tokens from Gemini-2.5 against a provider-reported target of 32948, while related Gemini and Claude models show substantial length and semantic alignment despite unavailable ground truth. The findings are relevant to providers including Google, Anthropic, and OpenAI: the study reports possible exposure of system-prompt content from OpenAI o4-mini and shows that a defensive system prompt lowers mean ASR@90 to 5.0%, but adaptive attacks still succeed. Removing reasoning state after tool calls is the strongest tested mitigation, indicating that protecting hidden CoT requires both state isolation and reduced API fidelity signals.
Original abstract
Hidden chain-of-thought (CoT) traces, especially those from frontier proprietary large reasoning models (LRMs), are valuable model assets. Yet whether these hidden CoTs can be directly extracted from black-box models remains largely unexplored. In this work, we systematically study whether hidden CoTs can be extracted near-verbatim from black-box LRMs through API interactions. We identify a previously overlooked reasoning replay surface between tool calls and develop EchoCoT, a multi-step attack that iteratively extracts hidden CoTs using API-returned fidelity signals. We further develop an LLM-based optimization framework that automatically searches for an effective universal injection trajectory across various datasets. We evaluate EchoCoT on three open-source and five frontier proprietary LRMs. On open-source LRMs, EchoCoT achieves up to 66.4\% near-verbatim extraction success, with the extracted trace length within 10\% of the target and at least 90\% of tokens exactly matching the target CoT. The same injection trajectory also generalizes to unseen datasets, achieving up to 80\% extraction success under the same criterion. For tested frontier proprietary LRMs, a substantial fraction of extracted CoTs closely align with provider-reported reasoning lengths and available CoT summaries. EchoCoT can also extract very long CoTs: on Gemini-2.5, it extracts 33,463 tokens from a 32,948-token target. These results establish hidden-CoT extraction as a practical security risk and highlight the need to better protect hidden CoT assets.
Read the original paperMore in AI Safety
Browse all 39 papers →Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng
Helpful AI agents may secretly work around safety rules to assist one another, creating rare but serious information-leakage risks that compound over repeated interactions.
Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.