🛰️ Daily AI Frontier
‹ back to 2026-09-16

Large Language Models Develop Belief State Geometry In-Context

Research LLMs & Foundation Models

Ranking

Overall 78
Content 95
Popularity 39

Observed public metrics from 1 member.

Representative image for Large Language Models Develop Belief State Geometry In-Context

Merged summary

TL;DR - A controlled study finds that LLM activations encode hidden Markov model belief states during in-context learning. Causal interventions suggest this geometry is functionally involved in prediction, supporting the view that LLMs approximate Bayesian inference over context-inferred models.

  • Belief states were linearly decoded from residual-stream activations across six open-source LLMs and 40 non-trivial HMMs.
  • Peak probe performance ranged from (R^2=0.83) to (0.99), appearing at layers ranging from early to late.
  • Patching and steering the identified subspace preserved prediction quality near that of unmodified models, while control interventions substantially degraded it.
  • The results extend activation-geometry findings from HMM-trained toy networks to production-scale LLMs.

Sources (1)

Large Language Models Develop Belief State Geometry In-Context

arXiv cs.LG Daniel Balcells, Andrew Jun Lee, Chirag Rastogi, Paul M. Riechers, Adam Shai, Xavier Poncini 2026-09-15 arXiv:2609.17376
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-21 14:19:24.018904 UTC

TL;DR - A controlled study finds that LLM activations encode hidden Markov model belief states during in-context learning. Causal interventions suggest this geometry is functionally involved in prediction, supporting the view that LLMs approximate Bayesian inference over context-inferred models.

  • Belief states were linearly decoded from residual-stream activations across six open-source LLMs and 40 non-trivial HMMs.
  • Peak probe performance ranged from (R^2=0.83) to (0.99), appearing at layers ranging from early to late.
  • Patching and steering the identified subspace preserved prediction quality near that of unmodified models, while control interventions substantially degraded it.
  • The results extend activation-geometry findings from HMM-trained toy networks to production-scale LLMs.
item →