Animation of an ethanol and an aspirin molecule vibrating, with arrows for the forces predicted on each atom by eight models. Below, each model’s energy error and force error fall together over training.

Eight models per molecule, each with a different energy/force weighting. Arrows show the forces the eight models predict on each atom; below, their energy and force errors fall together during training.

Density functional theory (DFT) is too expensive for high-throughput materials screening, so we approximate it with neural networks.

Those networks have two jobs at once: predict a structure’s energy, and predict the force on every atom. The two aren’t independent, because the force is the gradient of the energy, so one network has to get both right.

That leaves one setting to choose: how much to weight energy against forces in the loss. The field mostly treats this as a trade-off. Push on forces and you lose energy accuracy, and whole methods exist to tune the balance during training.

So we swept the setting end to end: eight models per molecule, on MD17 ethanol and aspirin, with the force weight running from 0 to 1.

On these two molecules there was no trade-off. Energy and force accuracy improved together across almost the entire range. Energy-only training gave the worst energy model for both molecules, and adding force supervision cut energy error by 65%.

That makes sense once you count. Each structure gives you one energy but 3N force components, where N is the number of atoms, so the forces do most of the work of learning the surface.

With thanks to Bruno Andreis at the Torr Vision Group, Oxford, for the conversations.