Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models
Ranking
Overall
76
Content
90
Popularity
42
Observed public metrics from 1 member.
Merged summary
TL;DR - Chain-of-thought models often either finish successfully or exhaust their token budget with poor results. Hidden-state probes show a modest early signal of non-convergence, potentially enabling early exits and adaptive compute allocation.
- Converged generations achieved 90.3% AIME accuracy versus 6.6% for non-converged generations; 62.0% converged overall.
- A linear probe on layer-20 activations at token 150 reached AUC 0.608 (±0.080, 5-fold CV).
- Activation probes outperformed token-entropy and repetition-based behavioral baselines.
- The sweep-level permutation result (p=0.063) did not meet conventional significance thresholds.
Sources (1)
Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - Chain-of-thought models often either finish successfully or exhaust their token budget with poor results. Hidden-state probes show a modest early signal of non-convergence, potentially enabling early exits and adaptive compute allocation.
- Converged generations achieved 90.3% AIME accuracy versus 6.6% for non-converged generations; 62.0% converged overall.
- A linear probe on layer-20 activations at token 150 reached AUC 0.608 (±0.080, 5-fold CV).
- Activation probes outperformed token-entropy and repetition-based behavioral baselines.
- The sweep-level permutation result (p=0.063) did not meet conventional significance thresholds.