🛰️ Daily AI Frontier
‹ back to 2026-09-21

World Modeling in Transformers

arXiv cs.AI LLMs & Foundation Models Pierre Beckmann, Matthieu Queloz, Andre Freitas 2026-09-18

TL;DR - A mechanistic study of TaxiGPT finds that behavioral navigation failures do not necessarily imply the absence of a coherent world model. The transformer learns an internal Manhattan map and navigation mechanisms, but interference between superposed features can prevent it from using them reliably.

  • TaxiGPT represents intersections and streets, tracks its position, and uses a goal-directed “compass” to navigate.
  • Causal interventions trace failures to interference between superposed intersection features, which disrupts localization.
  • “Affordance packing” groups intersections with identical legal moves, limiting the impact of localization errors.
  • Mechanistic indicators show that distinct world-modeling capacities emerge at different stages of training.

view merged work →