Research fellowships
38 projects, opening in batches. Each project comes with the idea, the early experiments and, where there is one, the proof. Fellows do the remaining academic work, finishing the paper or releasing the tool, as co-authors.
Our fellowships are for talented people from marginalised communities, wherever they are.
The terms
- Remote
- Work from anywhere.
- 1 to 12 months
- Most projects take three to six.
- Full compute
- We provide what the project needs.
- Co-authorship
- On the paper or the release you finish.
- Unpaid
- Fellowships don’t come with a salary or stipend.
How far along each project is
Every project has a readiness label, so you know what’s left before you start.
- Ready to finish (14)The results are in. The paper or the release is the remaining work.
- Decisive experiment first (12)One experiment decides the story, then comes the write-up.
- Starter project (10)A focused first project of one to three months.
- Negative-result write-up (2)A thorough negative result, ready to be written up.
The projects
One-line summaries for now. Each project gets a fuller page when its batch opens.
No projects match both filters.
Hallucination and verification
14 projects
-
Project 1
Checking claims with isolated provers
Make a claim-certification method’s error guarantee hold when the data shifts, test it on new kinds of claim, and release the benchmark.
-
Project 6
Teaching AI to say “I don’t know”
Compare a reasoning loop that only answers when its steps check out against simpler baselines such as majority voting, on more tasks and models.
-
Project 7
Can a small model spot its own reasoning mistakes?
Repeat step-level error detection on harder maths, on code and on other model families, with labels checked by people.
-
Project 8
When does more reasoning make answers worse?
Test a scaling law for how reasoning commits to an answer on larger models and real failures, and predict the most reliable reasoning length.
-
Project 10
When evidence checks go wrong
Turn stress tests of an open-source evidence checker into a catalogue of failure modes, and run the tests on other grounding tools.
-
Project 11
Is this text machine-written?
Build a held-out benchmark for spotting AI-written text across languages and domains, and compare with established detectors.
-
Project 15
Will AI agents actually check their work?
Design and compare ways to build checking into coding agents, on a public benchmark large enough to tell the designs apart.
-
Project 16
Independent checks for AI-written code
Build a benchmark of AI-written code with known bugs, and compare checks by running the code, by a second model and by formal methods.
-
Project 26
Proving an answer is grounded without revealing the evidence
Design a reliable check of what a model already knows, for zero-knowledge proofs that answers are grounded, then rerun the experiments and finish the paper.
-
Project 27
Certifying the evidence behind a bug diagnosis
Scale certified evidence selection for bug diagnoses to a few hundred real GitHub issues, and measure how often wrong explanations get certified.
-
Project 28
Checking each step of an agent’s reasoning
Run a step-grounding scorer on several hundred real agent steps, and compare it with word-overlap and entailment baselines.
-
Project 29
Why does a hallucination detector work on one model and not another?
Rebuild a hallucination benchmark without its shortcut, and map which model families a memory-based detector works on, and why.
-
Project 33
How much room does a text have for a hidden watermark?
Validate a key-free measure of how much a text could carry a hidden watermark, across watermarking schemes and models.
-
Project 36
Auditing published machine learning claims with an adversarial agent
Audit 10 to 20 published machine learning claims with an adversarial agent workflow, check the verdicts by hand, and publish a report.
In-context learning
6 projects
-
Project 9
Do helpful notes make problems easier?
Repeat two experiments on what helpful notes give a model, on open models of different sizes, and combine them into one paper.
-
Project 23
Do order effects survive on frontier models?
Reproduce the order-effect measurement from our ICML 2026 paper and test how strongly it holds across newer model sizes and families.
-
Project 24
Does low training loss predict what a model learns?
Test whether low training loss on a dataset predicts what a model learns from it, across many datasets, tasks and model sizes.
-
Project 25
How does a model count evidence in its prompt?
Replicate a mechanistic account of how models count evidence in a prompt, across model families and sizes.
-
Project 34
Auditing a forecasting model’s sensitivity to bad data
Find out when exact attention responses beat gradients for auditing forecasting models, across several forecasting datasets.
-
Project 35
Predicting when an in-context skill switches on
Test a law for when in-context skills switch on across more models, and explain why its slopes transfer but its switch-on points don’t.
Changing models without retraining
7 projects
-
Project 2
Order-proof few-shot answers
Show that a small controller gives order-averaged few-shot answers in one run across more tasks and model families, then finish the paper.
-
Project 4
A lighter alternative to fine-tuning
Run a properly powered comparison of learned gates against LoRA, IA3 and ReFT, and finish the paper from the existing drafts.
-
Project 30
Prompts as weight changes, or prompts as memory?
Test injecting a prompt’s stored attention memory directly, on more behaviours and models, and how far it can be compressed.
-
Project 31
Finding the shortest instruction that provably works
Run a certified search for the shortest instruction that works on more tasks and models, and compare it with prompt-optimisation methods.
-
Project 32
Personalising image models with gates instead of LoRA
Evaluate gate-based personalisation of image generators on a standard benchmark, and fix the case where a subject and a style are combined.
-
Project 37
Budgeted guidance for image generators: writing up a negative result
Write up a thorough negative result on information-budgeted guidance for image generators, and release the measurement code.
-
Project 38
What control at inference time can’t do: three negative results
Write one paper covering three negative results on changing model behaviour at inference time, and release the code.
Faster, cheaper AI
4 projects
-
Project 3
Compressing a model’s memory with a guarantee
Benchmark per-prompt compression of a model’s memory against standard methods on long-context suites, and measure real memory and speed savings.
-
Project 5
What must an AI remember, and what can it work out again?
Improve how we test what a model knows without its memory, run on a public memory benchmark, and define what the certificate guarantees.
-
Project 18
Better 4-bit models with exact rearrangements
Compare weight rearrangements chosen from the weights alone with trained methods such as FPTQuant and CafeQ, head to head.
-
Project 21
Shrinking a coding agent’s context to 5%
Rerun a context-compression method for coding agents on several hundred tasks, against simple baselines and existing compressors.
World models and science
4 projects
-
Project 12
Do students inherit their teacher’s invariances?
Find out why strong graph-network teachers don’t pass their symmetry on to student models, with proper controls, and write the paper.
-
Project 13
Why compression breaks symmetry
Turn a theorem on when limited models break symmetry into a paper: full proofs, extensions, and experiments in trained networks.
-
Project 17
Which biological signals does a model really use?
Extend a masked-prediction benchmark on cancer genomics data, and separate the relationships that survive negative controls from artefacts.
-
Project 22
Teaching models to ignore what shouldn’t matter
Find measurable properties that predict when averaging over nuisances helps, across tasks where the method works and where it fails.
Maths, cryptography and more
3 projects
-
Project 14
Erdős–Straus: a verified library and a map of what’s left
Release a Lean-verified library on the Erdős–Straus conjecture, finish the expository paper, and attack the cases that remain open.
-
Project 19
Fast proof-of-work kernels for matrix multiplication
Write the threat model for proof-of-work GPU kernels, attempt a soundness argument or find the attack, and benchmark at realistic sizes.
-
Project 20
Does a charge-and-discharge memory help attention recall?
Run the control that decides whether a new attention memory mechanism works, then test it on irregular-timing tasks if it holds up.
Who has worked with us
In 2025 we called for open contributions from underrepresented researchers around the world. Together we wrote papers on why language models hallucinate, and built open-source tools with more than 2,000 GitHub stars.
People who’ve researched with us have co-authored a paper at ICML 2026 and our 2026 preprints. Student researchers who’ve worked with us have gone on to PhDs at UCL and the University of Barcelona.
Questions
Read the open problems behind several projects, or email us.
For research questions, collaborations and press. For fellowships, the newsletter announces each batch first: lcphys.substack.com.