World models that respect physics
Can a model trained on data alone learn the symmetries of the world? Under a log-loss objective, breaking a symmetry is often the cheaper way to compress the data, so more data and compute don’t fix it. We measure where models break symmetries and distil students that don’t.
- Robot world models The same actions, written as joint targets or as deltas, produce different predictions. Our 2026 preprint with Ahmed Karim shows it, and proposes a symmetry-marginalised action space as the fix.
- Mezzanine Averages a teacher’s predictions over symmetry-equivalent inputs, then distils a student that gives that average in one pass. The spread across views, the “warrant gap”, measures how much the teacher was being swayed.
- Molecular dynamics in one pass A Lennard-Jones simulation that took 120,000 steps on an A100, distilled into a symmetry-stable model small enough for a phone, which recovers the same phase-level inference in one pass.
- A patch, not a cure You have to know the symmetry, and no general fix exists. Mezzanine helps when the variation that shouldn’t matter is real and compact, and we report where it doesn’t.
Read the paper → Mezzanine on GitHub →