World Modeling in Transformers
Ranking
Overall
81
Content
100
Popularity
36
Observed public metrics from 1 member.
Merged summary
TL;DR - A mechanistic study of TaxiGPT finds that behavioral navigation failures do not necessarily imply the absence of a coherent world model. The transformer learns an internal Manhattan map and navigation mechanisms, but interference between superposed features can prevent it from using them reliably.
- TaxiGPT represents intersections and streets, tracks its position, and uses a goal-directed “compass” to navigate.
- Causal interventions trace failures to interference between superposed intersection features, which disrupts localization.
- “Affordance packing” groups intersections with identical legal moves, limiting the impact of localization errors.
- Mechanistic indicators show that distinct world-modeling capacities emerge at different stages of training.
Sources (1)
World Modeling in Transformers
Public signals
Hugging Face upvotes 0 · Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - A mechanistic study of TaxiGPT finds that behavioral navigation failures do not necessarily imply the absence of a coherent world model. The transformer learns an internal Manhattan map and navigation mechanisms, but interference between superposed features can prevent it from using them reliably.
- TaxiGPT represents intersections and streets, tracks its position, and uses a goal-directed “compass” to navigate.
- Causal interventions trace failures to interference between superposed intersection features, which disrupts localization.
- “Affordance packing” groups intersections with identical legal moves, limiting the impact of localization errors.
- Mechanistic indicators show that distinct world-modeling capacities emerge at different stages of training.