Animation of an aeroplane made of points. It is split into patches, 60% of them are hidden, and a model learns to fill them back in as its error falls. The model’s embedding of a wing is then passed, with a written question, to a language model that answers in plain text.
Masked reconstruction is a simple idea: hide part of the input and train a model to rebuild it from what’s left. BERT does it with words, MAE with image patches and Point-MAE with 3D point clouds.
It works well on geometry, and I think symmetry is a big part of why. Real objects are rarely arbitrary. Animals that move forward tend to be bilaterally symmetric, and a lathe can only make round things. One part of a shape usually tells you a lot about another.
That’s exactly what reconstruction rewards. Hide a patch whose mirror image is still visible and the answer is already in the input. One of the best ways for the model to improve is to learn the symmetry and use it. With heavy masking, simple smoothing is no longer enough.
This is one way world models learn what real shapes look like, and it loosely mirrors human perception. We expect a smooth surface to stay smooth rather than break into a sudden spike, and we assume the hidden side of a mug matches the side we can see. A model trained this way picks up similar regularities from data.
Once trained, the model’s internal activations work as embeddings that summarise a shape’s structure. You can reuse them to classify parts, find similar designs or flag anomalies. The animation sketches one more option: aligning a wing’s embedding with text and passing it to a language model alongside a question, so it can answer in plain language.
The same prior can mislead. A model that expects symmetry can rebuild it where none exists. That matters most when you trust the filled-in parts and the odd shape is the point, such as a defect or a tumour.