Research fellowships

38 projects, opening in batches. Each project comes with the idea, the early experiments and, where there is one, the proof. Fellows do the remaining academic work, finishing the paper or releasing the tool, as co-authors.

Our fellowships are for talented people from marginalised communities, wherever they are.

The terms

Remote
Work from anywhere.
1 to 12 months
Most projects take three to six.
Full compute
We provide what the project needs.
Co-authorship
On the paper or the release you finish.
Unpaid
Fellowships don’t come with a salary or stipend.

We open projects in batches, and the newsletter announces each one first. Sign up at lcphys.substack.com to hear when the next batch opens, or browse the projects below.

How far along each project is

Every project has a readiness label, so you know what’s left before you start.

  • Ready to finish (14)The results are in. The paper or the release is the remaining work.
  • Decisive experiment first (12)One experiment decides the story, then comes the write-up.
  • Starter project (10)A focused first project of one to three months.
  • Negative-result write-up (2)A thorough negative result, ready to be written up.

The projects

One-line summaries for now. Each project gets a fuller page when its batch opens.

Hallucination and verification

14 projects

  • Project 1

    Checking claims with isolated provers

    Make a claim-certification method’s error guarantee hold when the data shifts, test it on new kinds of claim, and release the benchmark.

    Ready to finish 3–6 monthsMachine learningStatistics
  • Project 6

    Teaching AI to say “I don’t know”

    Compare a reasoning loop that only answers when its steps check out against simpler baselines such as majority voting, on more tasks and models.

    Ready to finish 3–6 monthsMachine learningNLP
  • Project 7

    Can a small model spot its own reasoning mistakes?

    Repeat step-level error detection on harder maths, on code and on other model families, with labels checked by people.

    Ready to finish 3–6 monthsMachine learningNLP
  • Project 8

    When does more reasoning make answers worse?

    Test a scaling law for how reasoning commits to an answer on larger models and real failures, and predict the most reliable reasoning length.

    Ready to finish 3–6 monthsMachine learningTheory and maths
  • Project 10

    When evidence checks go wrong

    Turn stress tests of an open-source evidence checker into a catalogue of failure modes, and run the tests on other grounding tools.

    Ready to finish 2–3 monthsMachine learningNLP
  • Project 11

    Is this text machine-written?

    Build a held-out benchmark for spotting AI-written text across languages and domains, and compare with established detectors.

    Ready to finish 3–6 monthsNLPMachine learning
  • Project 15

    Will AI agents actually check their work?

    Design and compare ways to build checking into coding agents, on a public benchmark large enough to tell the designs apart.

    Decisive experiment first 3–6 monthsML engineeringMachine learning
  • Project 16

    Independent checks for AI-written code

    Build a benchmark of AI-written code with known bugs, and compare checks by running the code, by a second model and by formal methods.

    Decisive experiment first 3–6 monthsML engineeringMachine learning
  • Project 26

    Proving an answer is grounded without revealing the evidence

    Design a reliable check of what a model already knows, for zero-knowledge proofs that answers are grounded, then rerun the experiments and finish the paper.

    Decisive experiment first 3–6 monthsMachine learningCryptography
  • Project 27

    Certifying the evidence behind a bug diagnosis

    Scale certified evidence selection for bug diagnoses to a few hundred real GitHub issues, and measure how often wrong explanations get certified.

    Starter project 1–3 monthsMachine learning
  • Project 28

    Checking each step of an agent’s reasoning

    Run a step-grounding scorer on several hundred real agent steps, and compare it with word-overlap and entailment baselines.

    Starter project 1–3 monthsML engineeringMachine learning
  • Project 29

    Why does a hallucination detector work on one model and not another?

    Rebuild a hallucination benchmark without its shortcut, and map which model families a memory-based detector works on, and why.

    Starter project 1–3 monthsInterpretabilityMachine learning
  • Project 33

    How much room does a text have for a hidden watermark?

    Validate a key-free measure of how much a text could carry a hidden watermark, across watermarking schemes and models.

    Starter project 1–3 monthsNLPStatistics
  • Project 36

    Auditing published machine learning claims with an adversarial agent

    Audit 10 to 20 published machine learning claims with an adversarial agent workflow, check the verdicts by hand, and publish a report.

    Starter project 2–3 monthsML engineeringMachine learning

In-context learning

6 projects

  • Project 9

    Do helpful notes make problems easier?

    Repeat two experiments on what helpful notes give a model, on open models of different sizes, and combine them into one paper.

    Ready to finish 3–6 monthsMachine learningStatistics
  • Project 23

    Do order effects survive on frontier models?

    Reproduce the order-effect measurement from our ICML 2026 paper and test how strongly it holds across newer model sizes and families.

    Decisive experiment first 3–6 monthsStatisticsMachine learning
  • Project 24

    Does low training loss predict what a model learns?

    Test whether low training loss on a dataset predicts what a model learns from it, across many datasets, tasks and model sizes.

    Decisive experiment first 2–4 monthsMachine learning
  • Project 25

    How does a model count evidence in its prompt?

    Replicate a mechanistic account of how models count evidence in a prompt, across model families and sizes.

    Decisive experiment first 3–6 monthsInterpretabilityMachine learning
  • Project 34

    Auditing a forecasting model’s sensitivity to bad data

    Find out when exact attention responses beat gradients for auditing forecasting models, across several forecasting datasets.

    Starter project 1–3 monthsTime seriesMachine learning
  • Project 35

    Predicting when an in-context skill switches on

    Test a law for when in-context skills switch on across more models, and explain why its slopes transfer but its switch-on points don’t.

    Starter project 1–3 monthsTheory and mathsMachine learning

Changing models without retraining

7 projects

  • Project 2

    Order-proof few-shot answers

    Show that a small controller gives order-averaged few-shot answers in one run across more tasks and model families, then finish the paper.

    Ready to finish 3–6 monthsMachine learningML engineering
  • Project 4

    A lighter alternative to fine-tuning

    Run a properly powered comparison of learned gates against LoRA, IA3 and ReFT, and finish the paper from the existing drafts.

    Ready to finish 3–6 monthsMachine learningML engineering
  • Project 30

    Prompts as weight changes, or prompts as memory?

    Test injecting a prompt’s stored attention memory directly, on more behaviours and models, and how far it can be compressed.

    Starter project 1–3 monthsMachine learning
  • Project 31

    Finding the shortest instruction that provably works

    Run a certified search for the shortest instruction that works on more tasks and models, and compare it with prompt-optimisation methods.

    Starter project 1–2 monthsMachine learningNLP
  • Project 32

    Personalising image models with gates instead of LoRA

    Evaluate gate-based personalisation of image generators on a standard benchmark, and fix the case where a subject and a style are combined.

    Starter project 1–3 monthsImage generationMachine learning
  • Project 37

    Budgeted guidance for image generators: writing up a negative result

    Write up a thorough negative result on information-budgeted guidance for image generators, and release the measurement code.

    Negative-result write-up 1–2 monthsImage generationStatistics
  • Project 38

    What control at inference time can’t do: three negative results

    Write one paper covering three negative results on changing model behaviour at inference time, and release the code.

    Negative-result write-up 1–2 monthsMachine learningStatistics

Faster, cheaper AI

4 projects

  • Project 3

    Compressing a model’s memory with a guarantee

    Benchmark per-prompt compression of a model’s memory against standard methods on long-context suites, and measure real memory and speed savings.

    Ready to finish 3–6 monthsML engineeringMachine learning
  • Project 5

    What must an AI remember, and what can it work out again?

    Improve how we test what a model knows without its memory, run on a public memory benchmark, and define what the certificate guarantees.

    Ready to finish 3–6 monthsMachine learningStatistics
  • Project 18

    Better 4-bit models with exact rearrangements

    Compare weight rearrangements chosen from the weights alone with trained methods such as FPTQuant and CafeQ, head to head.

    Decisive experiment first 3–6 monthsMachine learningML engineering
  • Project 21

    Shrinking a coding agent’s context to 5%

    Rerun a context-compression method for coding agents on several hundred tasks, against simple baselines and existing compressors.

    Decisive experiment first 2–4 monthsML engineering

World models and science

4 projects

  • Project 12

    Do students inherit their teacher’s invariances?

    Find out why strong graph-network teachers don’t pass their symmetry on to student models, with proper controls, and write the paper.

    Ready to finish 3–6 monthsMaterials scienceMachine learning
  • Project 13

    Why compression breaks symmetry

    Turn a theorem on when limited models break symmetry into a paper: full proofs, extensions, and experiments in trained networks.

    Ready to finish 3–6 monthsTheory and maths
  • Project 17

    Which biological signals does a model really use?

    Extend a masked-prediction benchmark on cancer genomics data, and separate the relationships that survive negative controls from artefacts.

    Decisive experiment first 6–12 monthsBiologyMachine learning
  • Project 22

    Teaching models to ignore what shouldn’t matter

    Find measurable properties that predict when averaging over nuisances helps, across tasks where the method works and where it fails.

    Decisive experiment first 3–6 monthsPhysicsMachine learning

Maths, cryptography and more

3 projects

  • Project 14

    Erdős–Straus: a verified library and a map of what’s left

    Release a Lean-verified library on the Erdős–Straus conjecture, finish the expository paper, and attack the cases that remain open.

    Ready to finish 3–6 monthsNumber theory and Lean
  • Project 19

    Fast proof-of-work kernels for matrix multiplication

    Write the threat model for proof-of-work GPU kernels, attempt a soundness argument or find the attack, and benchmark at realistic sizes.

    Decisive experiment first 3–6 monthsCryptographyGPU programming
  • Project 20

    Does a charge-and-discharge memory help attention recall?

    Run the control that decides whether a new attention memory mechanism works, then test it on irregular-timing tasks if it holds up.

    Decisive experiment first 2–4 monthsMachine learning

Who has worked with us

In 2025 we called for open contributions from underrepresented researchers around the world. Together we wrote papers on why language models hallucinate, and built open-source tools with more than 2,000 GitHub stars.

People who’ve researched with us have co-authored a paper at ICML 2026 and our 2026 preprints. Student researchers who’ve worked with us have gone on to PhDs at UCL and the University of Barcelona.

Questions

Read the open problems behind several projects, or email us.

For research questions, collaborations and press. For fellowships, the newsletter announces each batch first: lcphys.substack.com.