← Press & talks
Talk slides

Training-Free Self-Improving Agents

Improving a frozen model: an information-theoretic view.

Leon Chlon · Agentic AI Foundation London × Prolific: Post-Training, Evals & Self-Improving Agents · London, 30 September 2026

Once training ends, a model’s weights stop changing. What can it still do to get better? This talk counts it in bits. After training, only the context can bring in new information. A hallucination is an answer given on too few bits, and because no model can certify itself, the missing bits have to come from outside. The last part shows a model repairing how it reads its own KV cache with exact algebra: no new data, no new weights.

The animations play as you scroll. Use ← → to step through the slides.

  1. Title slide: Training-Free Self-Improving Agents. Improving a frozen model: an information-theoretic view. Leon Chlon.
    1Training-Free Self-Improving Agents
  2. Information is surprise, measured in bits. After “The cat sat on the …”, “mat” has probability 1 in 2 (1 bit) and “piano” 1 in 1,024 (10 bits). Without the context “mat” is 1 in 8,192, or 13 bits, so the context supplied 12. Definitions of surprisal, information gain and KL divergence.
    2Information is surprise, measured in bits
  3. LLMs autocomplete from weights or context. L(y | c) = L(y) − i(y ; c): what is still missing is what the weights leave open minus what the context supplies. Animated diagram of a model as a closed system of weights and KV cache. After training the weights are frozen, so only the context can still change.
    3LLMs autocomplete from weights or context
  4. Four ways a model can improve itself. Training on its own outputs changes the weights but brings no new bits without an outside reward. Rewriting its context or memory brings in new bits from the world, if it looks. Searching, sampling, voting and averaging add no new bits but make existing ones usable (averaging orderings: 42.0% to 45.2%). Steering how it reads re-weights what it already has. A closed loop learns nothing new about the world.
    4Four ways a model can improve itself
  5. Same evidence, new order, a different answer. Animated diagram: the same claim and four pieces of evidence, asked in different orders, get probabilities from 0.18 to 0.93. Order sensitivity grows like a + b log n with the number of evidence chunks. A hallucination is confidence the evidence can’t back.
    5Same evidence, new order, a different answer
  6. When the evidence runs short, it asks for help. Δ = L(y) − L(y | c): the bits the evidence supplies, found by asking with and without it. Animated diagram: with too few bits (ISR 0.44) the model says it doesn’t know and asks for more evidence; with enough it answers. It can’t create the missing bits, but it can measure the shortfall and ask for them.
    6When the evidence runs short, it asks for help
  7. How many outside bits does trust cost? B2T = KL(Ber(p⋆) ‖ Ber(q)) is the bits needed to move a belief q to reliability p⋆; ISR is the bits the evidence supplies over the bits needed. Answer if ISR is at least 1, otherwise fetch bits or abstain. In a held-out audit of 528 items, the gate at ISR = 1 gave 0.0–0.7% hallucination while abstaining on 20.6–27.9%.
    7How many outside bits does trust cost?
  8. No model can certify itself: Turing on halting (1936), Rice on program properties (1953) and Gödel on consistency (1931). Self-checks are incomplete or unsound, so trusted verdicts come from outside. With outside, isolated provers, false approvals fell from 67% to 1.8%. Each outside yes/no carries at most 1 bit.
    8No model can certify itself
  9. It can repair its own KV cache using exact algebra. s = q · R(Δ) k: RoPE turns each key by its position. Animated diagram: a fact 4,096 tokens back is still in the cache but goes unread; giving it the phase of a near token raises accuracy from 12.5% to 77.5%.
    9It can repair its own KV cache using exact algebra
  10. Using RoPE as a derivative the model can experiment on itself. Animated diagram: the model scores 109,080 candidate edits to how it reads its own prompt, re-weights the note it ignored, and passes 6 of 6 tests. 95–97% of 92,160 position edits were predicted in the right direction.
    10Using RoPE as a derivative the model can experiment on itself
  11. Summary, three things to keep. Rare information needs more bits, and after training only the context can supply them. Loops that only talk to themselves re-use bits; new bits come from outside, and ISR says how many. No model can certify itself, so spend outside verdicts where the budget runs short.
    11Three things to keep