Gaussian, Dirichlet and Beta processes were my favourite models in 2013, when I had no compute and little data. They’re machine learning’s version of necessity being the mother of invention, and I still think about problems in their terms.
Bayesian methods let you make up for having little data by supplying your best guess at what the function, features or clusters might look like. That guess is the prior, and you supply it alongside the data.
The Gaussian process, for functions. A Gaussian process combines the prior with observations and returns a distribution over functions that could have generated the data, with more weight on functions that your data or your prior support. A new observation pulls nearby functions towards it and leaves distant ones alone. The mean of the distribution is your forecast, and its variance shrinks as data arrive. Today Gaussian processes help decide which expensive experiments are worth running: where we know least, and where we’re likely to learn most.
The Dirichlet process, for clustering. Suppose you want to cluster your data but can’t guess how many groups there are. Give a Dirichlet process a prior over infinitely many groups and it returns a distribution over ways of clustering your observations. Most groups end up with weights so close to zero that they can be ignored; the ones that survive are the clusters the data support. A new observation joins a group in proportion to its size or, with some probability, starts a new one, so the number of groups grows with the data.
The Beta process, for features. Now suppose each observation is a bundle of traits: an image is edges and textures, a text is meaning and style. With few observations you don’t want to spend them choosing features by hand. A Beta process takes a prior over infinitely many features and returns a distribution over subsets of them, keeping only the features the data support. A new observation inherits popular features and, with some probability, introduces new ones.
The same instinct, letting the data decide how many clusters or features there are, runs through many modern methods, including sparse feature decompositions used to interpret neural networks.