When AI fails, and why
Hassana Labs is an independent non-profit AI research lab. We work out when models will fail and why, then turn the maths into open-source tools.
- ICML 2026
- Predictable Compression Failures, our paper on when a model should answer and when it should abstain.
- Open source
- Berry, our verifier for AI answers, was featured at NVIDIA GTC 2026. Our open-source tools have more than 2,000 GitHub stars between them.
- Press
- Quoted in WIRED in August 2026, on how the watermarks in AI-written text can be removed.
- Support
- Our work has been supported by Survival and Flourishing Corp, Microsoft and Google.
- Founder
- Leon Chlon, PhD in machine learning from Cambridge, with research posts at Harvard Medical School, MIT and Oxford.
Research
Five areas with one thread: working out when and why AI fails, with maths that predicts it and tools that check it.
-
Hallucination and verification
Why do models get facts wrong when the evidence is in front of them, and how can we catch it?
Our ICML 2026 paper treats hallucination on evidence-grounded yes-or-no questions as a measurable shortfall of information. When the evidence doesn’t carry enough to support an answer, the model fills the gap with something plausible. That gives a rule for when to answer and when to abstain.
In a pre-specified audit of 528 held-out questions, the rule kept hallucinations to 0.0–0.7% while abstaining on 20.6–27.9% of them (95% confidence intervals).
Where it stops working. The gate trades coverage for reliability. It abstained on a fifth to a quarter of questions, and in later tests it also dropped 10–20% of good claims.
Berry, our open-source verifier, puts the same idea to work. It checks each claim in an answer against the evidence cited for it and keeps a tamper-evident record of every decision, inside Claude Code, Cursor, Codex or Gemini CLI.
The answer-or-abstain gate from our ICML 2026 paper. ISR is the ratio of the bits the evidence supplies to the bits a trustworthy answer needs. Below 1 the model says it doesn’t know; at 1 or above it answers. -
In-context learning
What is a model actually computing when it learns from examples in a prompt?
Across 92,160 positional edits on held-out prompts, our exact formula for attention predicted the direction of the change 95.4–96.5% of the time (Huang et al., 2026).
-
Changing models without retraining
Can a running model be adapted without retraining it?
On Qwen2.5-7B, a 50,000-parameter controller came within a point of LoRA on GSM8K (71% against 72%), and lost 4.7 points on held-out code where LoRA lost 16.3 (ntkmirror).
-
Faster, cheaper AI
How much of a large model’s memory can be dropped within a set error budget?
Our add-on makes Triton, a widely used tool for writing fast GPU programs, work on NVIDIA’s GB10 (Blackwell) desktop hardware until official support arrives (triton-blackwell).
-
World models and science
When do models trained on data learn the world’s symmetries, and when do they break them?
Write the same robot trajectory as absolute joint targets instead of changes, and a world model’s retrieval degrades 2.6 to 13.4 times across three robot datasets (Karim and Chlon, 2026).
Recent papers
All free to read. The papers page has a plain-English line for each.
Robot World Models Are Not Invariant to How the Actions Are Written
Exact Finite Attention Responses From RoPE Derivatives
Predictable Compression Failures: Order Sensitivity and Information Budgeting for Evidence-Grounded Binary Adjudication
LLMs are Bayesian in Expectation, Not Realization
All papers. Our code is on the open-source page: Berry, Mezzanine, ntkmirror, triton-blackwell.
Explainers
Short illustrated pieces on the ideas behind the work, first posted on LinkedIn.
-
Why a CNN beats a vision transformer on small data
Every symmetry you build into an architecture is data you don’t have to collect.
-
What masked reconstruction teaches a model about shape
Hide part of a shape, ask a model to rebuild it, and symmetry becomes a shortcut worth learning.
-
Noether’s theorem on a spring
Emmy Noether proved that every symmetry comes with a quantity that never changes, and a mass on a spring makes the idea easy to see.
-
The 6-or-9 problem
Why world models trained to compress what they see can learn the wrong thing about the world.
Fellowships
38 research projects are waiting for fellows. 14 are ready to finish, 12 need one decisive experiment first, 10 are starter projects and 2 are negative results to write up. Each project comes with the idea, the early experiments and, where there is one, the proof. Fellows do the remaining academic work, finishing the paper or releasing the tool, as co-authors.
Our fellowships are for talented people from marginalised communities, wherever they are. Fellowships are remote, with full compute access and co-authorship. People who’ve researched with us have co-authored a paper at ICML 2026 and our 2026 preprints. Student researchers who’ve worked with us have gone on to PhDs at UCL and the University of Barcelona.
Browse the 38 projects, or get the newsletter, which announces each batch first.
Why Hassana
My grandma never went to high school, but she taught me that learning has no gates.
Hassana Labs is named after our founder’s grandmother. She loved learning, taught him mathematics, and made everyone around her feel valued. The lab carries that on. More about the lab