World models and science

Can a model trained on data alone learn the symmetries of the world, and where can careful, checkable AI make a difference in science?

Under a log-loss objective, breaking a symmetry is often the cheaper way to compress the data. Our analysis says more training data strengthens that incentive rather than removing it. We measure where models break symmetries, and distil students that don’t.

We test our methods on chemistry, materials and biology, where a confident wrong answer has a real cost.

Animation of an Ising model: a grid of spins, the same grid with every spin flipped, and a chart of how differently two models treat the pair over time. The teacher’s gap spikes close to 1; the student trained to respect the flip stays much lower.

An Ising model shown as itself and with every spin flipped, a change that shouldn’t alter the answer. The chart tracks how differently each model treats the pair: the teacher’s gap spikes close to 1, while the student trained to respect the flip stays much lower, though not at zero.

What we’ve found

  • Robot world models

    Hand a robot world model the same trajectory written as absolute joint targets instead of changes from the current state, and it collapses: retrieval degrades 2.6 to 13.4 times across three robot datasets, and goal-conditioned action selection falls from 53% to 15%. Averaging over both ways of writing the actions restores task performance, and a disagreement penalty lifts worst-case agreement from 0.78 to 0.995. Our preprint with Ahmed Karim also reports the dataset where averaging alone doesn’t help.

  • Mezzanine

    Averages a teacher model’s predictions over symmetry-equivalent inputs, then distils a student that gives that average in one pass. The spread across views, which we call the warrant gap, measures how much the teacher was being swayed.

  • Molecular dynamics in one pass

    A Lennard-Jones simulation that took 120,000 steps on an A100 GPU, distilled into a small model that gives a phase-level answer in one forward pass and is built not to change it when the particles are rotated or relabelled.

  • Neural surrogates for DFT

    Sweeping the weighting between energy and forces on MD17 ethanol and aspirin showed no trade-off across almost the entire range. Adding force supervision cut energy error by 65%.

  • Materials

    Band-gap predictions for a perovskite swung across symmetry-equivalent views of the same crystal. A student distilled on orbit-averaged labels was stable and more accurate.

  • Cancer multi-omics

    Benchmarks on tumour data from The Cancer Genome Atlas, where AI proposes constrained predictions for missing values that are checked against negative controls. On hold for now.

Animation of one perovskite crystal shown under twelve symmetry operations. The teacher model’s band-gap prediction jumps between about 0.5 and 2.9 eV; the distilled student’s stays between about 1.8 and 2.5 eV, close to the true 2.1 eV.

One perovskite crystal, shown to the model under twelve symmetry operations that leave it unchanged. The teacher’s band-gap prediction swings (standard deviation 0.77 eV); the student distilled on orbit-averaged labels barely moves (0.21 eV). The true value is 2.1 eV.

Animation of an ethanol and an aspirin molecule vibrating, with arrows for the forces predicted on each atom by eight models. Below, each model’s energy error and force error fall together over training.

Eight models per molecule, each with a different energy/force weighting. Arrows show the forces the eight models predict on each atom; below, their energy and force errors fall together during training.

Where it stops working

  • A patch, not a cure. You have to know the symmetry, and no general fix exists. Mezzanine helps when the variation that shouldn’t matter is real and compact, and we report where it doesn’t.

Distil the expectation, not a single view.

Projects

  • Teaching models to ignore what shouldn’t matter

    On hold

    A toolkit that averages a model’s predictions over changes that shouldn’t matter, then distils that into one fast model.

  • One model instead of eight in molecular simulation

    Ongoing

    Distilling an ensemble of eight chemistry models into one that’s as accurate and eight times cheaper to run.

  • Cancer multi-omics with guard-railed AI

    On hold

    Benchmarks where AI proposes constrained predictions for missing cancer data, checked against negative controls.