🛰️ Daily AI Frontier
‹ back to 2026-08-22

EchoCoT: Extracting Hidden Chain-of-Thought from Large Reasoning Models

Research LLM Security

Ranking

Overall 79
Content 95
Popularity 42

Observed public metrics from 1 member.

Merged summary

TL;DR - EchoCoT is a multi-step API attack that exploits reasoning replay around tool calls to extract hidden chain-of-thought traces from black-box large reasoning models. Its strong near-verbatim extraction results suggest hidden reasoning is a practical security and intellectual-property risk.

  • EchoCoT iteratively uses API-returned fidelity signals to recover hidden reasoning and automatically optimizes a universal injection trajectory.
  • Across three open-source models, it achieved up to 66.4% near-verbatim extraction success under strict length and token-match criteria.
  • The attack generalized to unseen datasets with up to 80% extraction success.
  • Tests on five proprietary models produced traces aligned with reported reasoning lengths and summaries, including a 33,463-token extraction from Gemini-2.5.

Sources (1)

EchoCoT: Extracting Hidden Chain-of-Thought from Large Reasoning Models

arXiv cs.CR Yiting Qu, Ziqing Yang, Chi Cui, Ye Leng, Junjie Chu, Yang Zhang 2026-08-20 arXiv:2608.20055
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-03 14:21:18.769127 UTC

TL;DR - EchoCoT is a multi-step API attack that exploits reasoning replay around tool calls to extract hidden chain-of-thought traces from black-box large reasoning models. Its strong near-verbatim extraction results suggest hidden reasoning is a practical security and intellectual-property risk.

  • EchoCoT iteratively uses API-returned fidelity signals to recover hidden reasoning and automatically optimizes a universal injection trajectory.
  • Across three open-source models, it achieved up to 66.4% near-verbatim extraction success under strict length and token-match criteria.
  • The attack generalized to unseen datasets with up to 80% extraction success.
  • Tests on five proprietary models produced traces aligned with reported reasoning lengths and summaries, including a 33,463-token extraction from Gemini-2.5.
item →