One of my favourite scientists is Ray Solomonoff, whose work on algorithmic induction in the early 1960s was one of the first theories of machine learning.
He started from Occam’s razor: the simplest explanation is usually the best. Then he made it precise for machines. To predict what comes next in a sequence of observations, consider every computer program that could have produced the data so far, measure each program’s length in bits, give shorter programs more weight, and update with Bayes’ rule as new data arrive.
Bayesian inference has critics because choosing the prior is a modelling decision. Even in linear regression you can choose a Gaussian or a Laplace prior to express how cautious you are about the coefficients, and people will argue with you either way.
Next time you’re deep in a hard problem, your models don’t make sense and the temptation is to add complexity and parameters, remember Solomonoff: smaller and simpler is usually better.