Changing models without retraining

Can we change how a model behaves while it runs, safely and without costly retraining?

Fine-tuning is expensive, and it can make a model forget what it knew. We showed that a training step on a pretrained transformer has a forward-pass counterpart: to first order, its effect equals a controlled change to the model’s internal signals. So a frozen model can be adapted by a small controller that never touches its weights.

A frozen model edits how it reads its prompt Three steps. One: the model scores 109,080 candidate edits to how it reads its own prompt. Two: it applies the top-ranked edit, which re-weights a note it had ignored. Three: it runs the tests again and all six pass. 1 Score 109,080 candidate edits to how the model reads its own prompt 2 Apply the top-ranked edit re-weight a note it had ignored 3 Run the tests again 6 of 6 pass
The worked example from our talk at the Agentic AI Foundation in London, as a diagram. Nothing in the model’s weights changes.

What we’ve found

  • ntkmirror

    On Qwen2.5-7B, a 50,000-parameter controller came within a point of LoRA’s accuracy on GSM8K, a benchmark of grade-school maths problems (71% against 72%), after 22 seconds of fitting. On held-out code generation it lost 4.7 points, where LoRA lost 16.3.

  • Controllers that add up

    Two controllers fitted separately combine by simple addition and keep each task’s gains. Merging two LoRA adapters the same way cost 17% on GSM8K.

  • A frozen model that edits how it reads

    In a worked example from our London talk, a frozen model scored 109,080 candidate edits to how it reads its own prompt, picked the one that re-weighted a note it had ignored, and then passed all six tests.

  • Order-proof few-shot answers

    A small controller gives, in one ordinary run, the answer you’d get by averaging over all orderings of the examples.

Where it stops working

  • We tested whether an information budget could steer image generators better than the standard method. It couldn’t, and we’re writing that up as a negative result.

Projects

  • Steering attention with exact predictions

    Ongoing

    A tool that ranks which parts of a prompt drive an answer, and predicts the effect of attention edits before making them.

  • Order-proof few-shot answers

    Completed

    A small controller that gives, in one ordinary run, the answer you’d get by averaging over all orderings of the examples.

  • A lighter alternative to fine-tuning

    On hold

    Adjusting a frozen model’s internal signals with small learned gates instead of retraining its weights.

  • Budgeted guidance for image generators

    On hold

    Testing whether an information budget can steer image generators better than the standard method. It couldn’t.