Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning
Ranking
Overall
83
Content
100
Popularity
42
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper tests whether the apparent clarity of chain-of-thought traces reveals which reasoning steps actually affect model performance. LLM judges identify important steps better than a prevalence baseline but remain far below the estimated noise ceiling, challenging their use for interpretability and process supervision.
- Step importance is defined as “advantage”: the change in expected reward caused by including a step, estimated through Monte Carlo rollouts.
- Capable LLM judges can detect some high-advantage steps, but much of their functional importance is not recoverable from the trace text.
- Fine-tuning a step-level critic substantially improves judgments for incorrect responses, while performance on correct responses remains far from the ceiling.
- The results caution against equating readable reasoning traces with faithful explanations, particularly in process reward modeling and generative critique.
Sources (1)
Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning
Public signals
Hugging Face upvotes 0
TL;DR - This paper tests whether the apparent clarity of chain-of-thought traces reveals which reasoning steps actually affect model performance. LLM judges identify important steps better than a prevalence baseline but remain far below the estimated noise ceiling, challenging their use for interpretability and process supervision.
- Step importance is defined as “advantage”: the change in expected reward caused by including a step, estimated through Monte Carlo rollouts.
- Capable LLM judges can detect some high-advantage steps, but much of their functional importance is not recoverable from the trace text.
- Fine-tuning a step-level critic substantially improves judgments for incorrect responses, while performance on correct responses remains far from the ceiling.
- The results caution against equating readable reasoning traces with faithful explanations, particularly in process reward modeling and generative critique.