Planting a Latent Variable in Natural-Looking Text: a More Realistic Test of Belief States in LLMs and Their Link to Concept Geometry
TL;DR - Researchers embed a controllable latent variable into natural-looking text and show that a small transformer trained on it tracks the variable’s Bayesian posterior. The model also reproduces the latent process’s ring geometry, linking learned concept geometry to underlying statistical dynamics.
- An LLM teacher generates ordinary text while being steered along one of eight unrelated sparse-autoencoder directions at each token.
- The active directions transition according to a ring-shaped Markov chain, creating a controlled latent process within realistic-looking language.
- A small transformer trained on the corpus learns to represent posterior beliefs about the active latent state.
- Its representations arrange the eight states in the Markov chain’s exact ring order, suggesting that latent-variable dynamics can shape feature geometry.