Research

Five areas, one thread: working out when and why AI fails, with maths that predicts it and tools that check it. Each area page says what we’ve found and where it stops working.

  • Why AI makes things up, and how to check it

    Why do models get facts wrong even when the evidence is in front of them, and can we catch it automatically?

    • 4 findings so far
    • 1 paper
    • 1 open-source tool
  • How models learn from context

    What is a model really computing when it learns from examples in a prompt, and can maths predict when it will fail?

    • 4 findings so far
    • 2 papers
  • Changing models without retraining

    Can we change how a model behaves while it runs, safely and without costly retraining?

    • 4 findings so far
    • 1 paper
    • 1 open-source tool
  • Faster, cheaper AI

    How can large models run with less memory and compute, with a guarantee on what we give up?

    • 3 findings so far
    • 1 open-source tool
  • World models and science

    Can a model trained on data alone learn the symmetries of the world, and where can careful, checkable AI make a difference in science?

    • 6 findings so far
    • 1 paper
    • 1 open-source tool

Open problems

Questions we think matter and haven’t answered.

  • Will AI agents actually check their work?

    Our verifier catches unsupported claims well on benchmarks. In one coding-agent trial, the agent never used it unless it was required to. How should checking be built into agent workflows so it happens reliably, without slowing everything down?

  • When does more reasoning make answers worse?

    Longer chains of thought can increase competition between candidate answers at the final step. Can we predict, for a given problem, the reasoning length that gives the most reliable answer?

  • Prompt injection as a family, not a string

    Real attacks vary in how they’re wrapped, where they appear and how they’re encoded. How can we give statistical guarantees about an AI system’s failure rate across a whole family of attack variants?

  • Independent checks for AI-written code

    A model can’t reliably check code it wrote itself, because it shares the blind spots that caused the bug. What’s the cheapest independent check that still catches most errors: running the code, a different model, or formal methods?

  • Which biological signals does a model really use?

    By hiding parts of a tumour’s multi-omics profile and asking a model to reconstruct them, we can see which relationships between data types it relies on. Which of those are reliable, and do they point to real biology?

  • What must an AI remember, and what can it work out again?

    Early results suggest an AI system’s store of facts can be trimmed a lot for questions it can reliably work out again, but recall of other facts suffers. Where is the right line, and can we certify it?