Large Language Models Develop Belief State Geometry In-Context
TL;DR - A controlled study finds that LLM activations encode hidden Markov model belief states during in-context learning. Causal interventions suggest this geometry is functionally involved in prediction, supporting the view that LLMs approximate Bayesian inference over context-inferred models.
- Belief states were linearly decoded from residual-stream activations across six open-source LLMs and 40 non-trivial HMMs.
- Peak probe performance ranged from (R^2=0.83) to (0.99), appearing at layers ranging from early to late.
- Patching and steering the identified subspace preserved prediction quality near that of unmodified models, while control interventions substantially degraded it.
- The results extend activation-geometry findings from HMM-trained toy networks to production-scale LLMs.