Hallucination and verification

Why AI makes things up, and how to check it

Why do models get facts wrong even when the evidence is in front of them, and can we catch it automatically?

Our ICML 2026 paper treats hallucination on evidence-grounded yes-or-no questions as a measurable shortfall of information. When the evidence doesn’t carry enough to support an answer, the model fills the gap with something plausible. That gives a rule for when a model should answer and when it should abstain.

We turn the same idea into open-source checks that test each claim against the evidence cited for it.

The answer-or-abstain gate A scale of ISR from 0 to 2 with a threshold at 1. Below 1 the model abstains, says it doesn’t know and asks for more evidence. At 1 or above it answers. An example from our talk sits at ISR 0.44, in the abstain region. Abstain says it doesn’t know, asks for more evidence Answer the evidence carries enough bits ISR = 1 0 1 2 ISR Example from our talk: ISR 0.44
The answer-or-abstain gate from the ICML paper. ISR is the ratio of the bits the evidence supplies to the bits a trustworthy answer needs. Below 1, the model says it doesn’t know and asks for more evidence; at 1 or above, it answers.

What we’ve found

  • Predictable Compression Failures

    Published at ICML 2026. In a pre-specified audit of 528 held-out questions, the paper’s answer-or-abstain gate kept hallucinations to 0.0–0.7% while abstaining on 20.6–27.9% of them (95% confidence intervals). The paper doesn’t claim that calibration holds for every model family or for open-ended writing.

  • Berry

    Our open-source verifier, formerly called Strawberry. It checks each claim against the evidence cited for it and keeps a tamper-evident record of every decision, inside Claude Code, Cursor, Codex or Gemini CLI.

  • Knowing when to stay quiet

    An abstention gate built on the paper raised the share of grounded claims among those a model shipped from about 50% to between 75% and 90%, depending on the model.

  • Checking with isolated provers

    A model can’t certify its own work, so trusted verdicts have to come from outside it. When claims were split into parts and checked by separate, isolated models, false approvals fell from 67% to 1.8% in our tests.

Where it stops working

  • The gate trades coverage for reliability. In the ICML audit it abstained on a fifth to a quarter of questions, and in later tests it also dropped 10–20% of good claims.
  • The order effects in the paper are measured on specific models, and we haven’t yet shown how strongly they hold on the newest ones.

Projects

  • Berry: flagging claims that aren’t backed by evidence

    Ongoing

    An open-source verifier that checks each claim against the evidence cited for it, in a single pass.

  • Hallucinations are predictable

    Completed

    Our ICML 2026 paper: hallucination as a measurable shortfall of information, with a rule for when to answer or abstain.

  • Checking claims with isolated provers

    Completed

    Splitting a claim into parts checked by separate, isolated models, so copied or misattributed evidence can’t slip through.

  • Can a small model spot its own reasoning mistakes?

    Completed

    Testing whether a small model can find the wrong steps in its own maths reasoning, and whether checking steps improves answers.

  • Teaching AI to say “I don’t know”

    On hold

    A step-by-step reasoning loop that commits only to answers it can check, and abstains otherwise.

  • Is this text machine-written?

    Ongoing

    A detector for AI-generated text that is most accurate when it can compare a response with its human-written source.