World Modeling in Transformers

MATS Fellow:

Pierre Beckmann

Authors:

Pierre Beckmann, Matthieu Queloz, Andre Freitas

Citations

Citations

Abstract:

Behavioral failures can make a transformer appear to lack a world model even when it has learned faithful representations of its environment. We demonstrate this in TaxiGPT, a transformer trained on random walks through Manhattan whose failures have been interpreted as evidence of an incoherent internal map. Through mechanistic analysis and causal interventions, we show that the model represents intersections and streets, tracks its position, and uses a goal compass to navigate. We trace its failures to interference between superposed intersection features, which disrupts localization within the internal map. Affordance packing, which groups representations of intersections with the same legal moves, helps limit the consequences of these errors. Finally, we propose mechanistic indicators that we use to compare models and show that world-modeling capacities emerge at different stages of training. Our findings motivate a shift from asking whether a model has a world model to mechanistically studying its world modeling: the interacting capacities through which it represents its environment and uses those representations to guide behavior.

Recent research

World Modeling in Transformers

Authors:

Pierre Beckmann

Date:

September 18, 2026

Citations:

You Are What You Read: Misalignment via In-Context Persona Induction

Authors:

Kyuhee Kim, Benjamin Berczi

Date:

September 6, 2026

Citations:

Frequently asked questions

What is the MATS Program?
How long does the program last?