Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models
TL;DR - Chain-of-thought models often either finish successfully or exhaust their token budget with poor results. Hidden-state probes show a modest early signal of non-convergence, potentially enabling early exits and adaptive compute allocation.
- Converged generations achieved 90.3% AIME accuracy versus 6.6% for non-converged generations; 62.0% converged overall.
- A linear probe on layer-20 activations at token 150 reached AUC 0.608 (±0.080, 5-fold CV).
- Activation probes outperformed token-entropy and repetition-based behavioral baselines.
- The sweep-level permutation result (p=0.063) did not meet conventional significance thresholds.