In-context learning

How models learn from context

What is a model really computing when it learns from examples in a prompt, and can maths predict when it will fail?

Averaged over the order of the examples, language models come close to ideal Bayesian reasoning. Any single ordering can drift from it, and we’ve traced the drift to the positional encoding.

Same evidence, different orders A claim and four pieces of evidence, E1 to E4, presented in different orders. Across the orders, the model’s probability that the claim is supported ranged from 0.18 to 0.93. Order 1 Claim E1 E2 E3 E4 Order 2 Claim E3 E1 E4 E2 … and other orders of the same four Probability the claim is supported 0 0.5 1 0.18 0.93
The same claim and the same four pieces of evidence, asked in different orders. The model’s probability that the claim is supported ranged from 0.18 to 0.93. A transformer without positional encoding gives the same answer for every order.

What we’ve found

  • Bayesian in expectation

    On Qwen2.5 7B and 14B, the gap to an ideal Bayesian reasoner is a few hundredths of a bit per prediction. A transformer trained without positional encoding is exactly order-blind; adding one back raises its order sensitivity by eight orders of magnitude.

  • Exact attention responses

    A closed-form account of how attention with rotary position encoding (RoPE) responds when part of the input is moved, removed or changed. Across 92,160 positional edits on held-out prompts, it predicted the direction of the change correctly 95.4–96.5% of the time.

  • A ledger for every token

    Tracing an answer token by token: how much each token in the model’s memory shifted or corrected the output. Next, we’re turning it into a bound for in-context learning.

  • Showing examples is not training

    We tested the popular idea that learning from examples in a prompt is the same as one step of training, and in the settings we tested it doesn’t hold.

Animation titled How RoPE hands the model a derivative. Under the words the cat sat on the mat, a wave shows how much weight the model gives the word cat as it reads on. A reader dot slides along the wave and loses its grip on cat; the slope under it, in green, predicts the next value before the next word arrives.

How the attention a model pays to an earlier word rises and falls as it reads on. With rotary position encoding the pattern is fixed, so the slope tells the model how that attention will change before the next word arrives. The exact-attention paper carries this through a whole finite step.

Where it stops working

  • The Bayesian result holds on average over orderings. Any single run of a prompt can still drift from it.

Projects

  • Are language models Bayesian?

    Ongoing

    Averaged over the order of their examples, models come close to ideal Bayesian reasoning; any single ordering can drift from it.

  • Predicting the effect of editing a prompt, exactly

    Ongoing

    An exact formula for how attention responds when parts of a model’s input are moved, removed or changed.

  • Is showing examples the same as training?

    Completed

    Testing the popular idea that learning from prompt examples equals a step of training, and finding it doesn’t.

  • Do helpful notes make problems easier?

    Completed

    Measuring whether notes help a model because of the information they carry, or for another reason.