Animation in two parts. A convolutional network slides one ear-detecting filter over a cartoon cat and gets the same score, 3.40, wherever the cat sits. A vision transformer trained on the cat in one corner fails to recognise it elsewhere, until it has been trained on the cat in six different places.
Train a convolutional neural network (CNN) and a vision transformer (ViT) from scratch on the same small dataset and the CNN usually wins, even though it’s the older design.
The reason is a symmetry. When you have too little data for the space you’re learning in, you need an assumption that shrinks the space. For images that assumption is translation invariance: a cat is a cat whether it’s in the top left of the photo or the bottom right.
A 224 by 224 colour photo is about 150,000 numbers. Learning anything in a space that big with no assumptions would take an enormous amount of data.
A convolution hard-codes the assumption. One small filter slides over the whole image with the same weights everywhere. If it learns to spot an ear in one corner, it can spot it in every corner. Move the cat and the feature map moves with it; take the maximum and you get the same answer.
The savings are large. AlexNet’s first layer, built as a fully connected layer, would need about 44 billion weights. Looking only at local 11 by 11 patches brings that down to about 105 million. Sharing the weights across positions brings it down to about 35 thousand.
Vision transformers get none of this for free. They cut the image into patches, tag each patch with its position, and have to learn from data that position doesn’t matter. Given enough data (the original results used 14 million to 300 million images) they catch up with CNNs and then overtake them. On a small dataset there aren’t enough examples to work it out.
Every symmetry you build into an architecture is data you don’t have to collect, and compute you don’t have to spend.