Shorter programs get more weight Three programs that all reproduce the same data, 3, 5 and 8 bits long. Each gets prior weight proportional to 2 to the power of minus its length, so the 3-bit program gets about 78% of the weight, the 5-bit one about 20% and the 8-bit one about 2%. length in bits share of prior weight Program A 78% Program B 20% Program C 2% All three reproduce the data. Weight ∝ 2−length
Three programs that all reproduce the same data, 3, 5 and 8 bits long. Weighting each by 2 to the power of minus its length gives the shortest about 78% of the prior.

One of my favourite scientists is Ray Solomonoff, whose work on algorithmic induction in the early 1960s was one of the first theories of machine learning.

He started from Occam’s razor: the simplest explanation is usually the best. Then he made it precise for machines. To predict what comes next in a sequence of observations, consider every computer program that could have produced the data so far, measure each program’s length in bits, give shorter programs more weight, and update with Bayes’ rule as new data arrive.

Bayesian inference has critics because choosing the prior is a modelling decision. Even in linear regression you can choose a Gaussian or a Laplace prior to express how cautious you are about the coefficients, and people will argue with you either way.

Next time you’re deep in a hard problem, your models don’t make sense and the temptation is to add complexity and parameters, remember Solomonoff: smaller and simpler is usually better.