When Steering Fails in Latent Reasoning: A Latent-to-Language Transition Gap
Ranking
Overall
79
Content
95
Popularity
43
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper finds that activation steering in latent chain-of-thought has much less influence on generated language than comparable steering of explicit reasoning. The results point to a latent-to-language transition gap that future steering methods must address.
- Task information remains identifiable within continuous latent thoughts, suggesting the weak control is not caused by its absence.
- Model output distributions change abruptly at the boundary between latent reasoning and language generation.
- Task-related directions provide substantially weaker bidirectional control in latent CoT than in explicit CoT.
- The latent-to-language interface emerges as a key target for evaluating and designing latent-steering techniques.
Sources (1)
When Steering Fails in Latent Reasoning: A Latent-to-Language Transition Gap
Public signals
Hugging Face upvotes 0
TL;DR - This paper finds that activation steering in latent chain-of-thought has much less influence on generated language than comparable steering of explicit reasoning. The results point to a latent-to-language transition gap that future steering methods must address.
- Task information remains identifiable within continuous latent thoughts, suggesting the weak control is not caused by its absence.
- Model output distributions change abruptly at the boundary between latent reasoning and language generation.
- Task-related directions provide substantially weaker bidirectional control in latent CoT than in explicit CoT.
- The latent-to-language interface emerges as a key target for evaluating and designing latent-steering techniques.