A 6 or a 9 A digit drawn on a square of floor. A person standing at the bottom edge reads it as 6. A person at the top edge, facing the other way, reads the same shape as 9. 6 reads 6 reads 9
One shape on the floor. Stand at the bottom and it reads 6; walk round to the top and it reads 9.

Draw the number 6 on the floor. Stand on opposite sides and one of you reads 6, the other 9. Physics has an answer to disagreements like this: look for the quantity that doesn’t depend on where you stand.

A world model trained under log loss from one side of the room will confidently say 6. Train it from the other side and it says 9. Train it from both sides, with its position given as input, and it learns “if standing here, 6; if standing there, 9”. That is a cheaper way to compress the data than learning that the shape is ambiguous and needs more context.

The symmetry-breaking representation wins on log loss because it’s cheaper than representing the genuine invariance. The optimiser is doing exactly what you asked it to.

The uncomfortable implication is that world models trained this way aren’t learning the world’s symmetries. They learn whatever compressed representation minimises code length for the architecture they have. Much of the argument that scaling will fix this assumes that enough data and compute will recover the true invariances.

In our 2026 preprint, Ahmed Karim shows a version of this in robot world models. Write the same commanded trajectory as absolute joint targets or as changes from the current state, and a world model trained on one way predicts a different future from the other. Averaging over both ways of writing the actions repairs most of the damage, and the paper reports where it doesn’t.

Read the paper