Projects will be theoretical in flavour, but I expect to work with fellows over the first few weeks to find projects that match their interests. They will generally center around the mechanistic estimation of quantities traditionally estimated via randomised methods (see for example https://arxiv.org/abs/2605.05179), or developing heuristic explanations for mathematical statements. Alternatively, fellows may take a more abstract route and work on the theory of heuristic arguments themselves, for example by broadening our understanding of No-coincidence Principles.
George Robinson is an independent researcher formerly at the Alignment Research Center (ARC), working on a systematic and theoretically grounded approach to mechanistic interpretability. He is now looking to lead a research effort in London supporting this agenda. Previously, he was a PhD student at Oxford University specialising in Algebraic Number Theory. He lives in London, and is a member of the London Initiative for Safe AI (LISA).
Essential:
​
Preferred:
​
Optional extras:
I will discuss a few projects with each fellow, and give time for fellows to work on each problem to get a sense of which ones they would like to move forward with. I would expect fellows to be working almost entirely on a single problem by the third week.
The Winter 2026 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.