NTH

EchoCoT: Extracting Hidden Chain-of-Thought from Large Reasoning Models

AuthorsYiting Qu, Ziqing Yang, Chi Cui, Ye Leng, Junjie Chu, Yang Zhang

August 30, 2026 2 min read
Watch on YouTube
The one-line take

EchoCoT shows that hidden reasoning traces from powerful AI models may be recoverable through carefully designed API interactions, creating a major security concern.

Key results

66.4%
Open-source ASR@90

Maximum near-verbatim extraction success achieved by EchoCoT-LTGO across DeepSeek-V4-Flash, Qwen3.5-Plus, and GLM-5.2.

80%
Unseen-dataset ASR@90

Maximum transfer success on MATH500, JEEBench, and LiveCodeBench.

33463
Gemini-2.5 extracted CoT

Tokens extracted by EchoCoT from Gemini-2.5.

32948
Gemini-2.5 target CoT

Provider-reported target reasoning-token count for the Gemini-2.5 example.

5.0%
Defensive system prompt ASR@90

Mean attack success rate after applying the tested defensive system prompt.

What the paper found

EchoCoT demonstrates that hidden chain-of-thought can be recovered from black-box reasoning models through a tool-call replay surface: while ordinary multi-turn interactions discard internal reasoning, tool calls preserve it within the same turn. The attack repeatedly asks a scratchpad tool to archive the model’s reasoning, evaluates each candidate using API-reported reasoning-token counts and compressed CoT summaries, and adapts subsequent injections. Its LLM-based INJECT-REFLECT-DISTILL optimizer searches universal trajectories using length-guided optimization, or LTGO when textual summaries are available. Across DeepSeek-V4-Flash, Qwen3.5-Plus, and GLM-5.2, EchoCoT-LTGO reaches up to 66.4% ASR@90, requiring both length error no greater than 10% and Token-EM of at least 0.90, and transfers to MATH500, JEEBench, and LiveCodeBench with success rates up to 80%. The method also produces long traces: it extracts 33463 tokens from Gemini-2.5 against a provider-reported target of 32948, while related Gemini and Claude models show substantial length and semantic alignment despite unavailable ground truth. The findings are relevant to providers including Google, Anthropic, and OpenAI: the study reports possible exposure of system-prompt content from OpenAI o4-mini and shows that a defensive system prompt lowers mean ASR@90 to 5.0%, but adaptive attacks still succeed. Removing reasoning state after tool calls is the strongest tested mitigation, indicating that protecting hidden CoT requires both state isolation and reduced API fidelity signals.

Original abstract

Hidden chain-of-thought (CoT) traces, especially those from frontier proprietary large reasoning models (LRMs), are valuable model assets. Yet whether these hidden CoTs can be directly extracted from black-box models remains largely unexplored. In this work, we systematically study whether hidden CoTs can be extracted near-verbatim from black-box LRMs through API interactions. We identify a previously overlooked reasoning replay surface between tool calls and develop EchoCoT, a multi-step attack that iteratively extracts hidden CoTs using API-returned fidelity signals. We further develop an LLM-based optimization framework that automatically searches for an effective universal injection trajectory across various datasets. We evaluate EchoCoT on three open-source and five frontier proprietary LRMs. On open-source LRMs, EchoCoT achieves up to 66.4\% near-verbatim extraction success, with the extracted trace length within 10\% of the target and at least 90\% of tokens exactly matching the target CoT. The same injection trajectory also generalizes to unseen datasets, achieving up to 80\% extraction success under the same criterion. For tested frontier proprietary LRMs, a substantial fraction of extracted CoTs closely align with provider-reported reasoning lengths and available CoT summaries. EchoCoT can also extract very long CoTs: on Gemini-2.5, it extracts 33,463 tokens from a 32,948-token target. These results establish hidden-CoT extraction as a practical security risk and highlight the need to better protect hidden CoT assets.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis