EchoCoT: Extracting Hidden Chain-of-Thought from Large Reasoning Models
TL;DR - EchoCoT is a multi-step API attack that exploits reasoning replay around tool calls to extract hidden chain-of-thought traces from black-box large reasoning models. Its strong near-verbatim extraction results suggest hidden reasoning is a practical security and intellectual-property risk.
- EchoCoT iteratively uses API-returned fidelity signals to recover hidden reasoning and automatically optimizes a universal injection trajectory.
- Across three open-source models, it achieved up to 66.4% near-verbatim extraction success under strict length and token-match criteria.
- The attack generalized to unseen datasets with up to 80% extraction success.
- Tests on five proprietary models produced traces aligned with reported reasoning lengths and summaries, including a 33,463-token extraction from Gemini-2.5.