Methodologically speaking, a proposed AI safety technique is only as good as its safety argument. This does not imply that we need to mathematically prove absolute safety from no assumptions (which is impossible), but it does mean we need to track what assumptions we need to make and rigorously connect them into an argument. The best (most real) risk arguments (eg "If Anyone Builds It, Everyone Dies") and safety arguments (eg basin-of-corrigibility arguments or alignment-by-default via natural abstractions) are still highly informal; I aim to come up with formal arguments, accepting some simplifications at first for the sake of progress, and moving towards increasingly realistic models.
My main line of research tackles this problem head-on by analyzing agentic trust directly, with the aim of clarifying how humans can justifiably trust AI systems. As of this writing, my current thinking about research directions for this project is described here. An older, more out-of-date but more thorough description of research directions is here. If this line of research goes well, it would produce relevant advice for AI engineering (such as LLM work).
Historically, I have focused on somewhat esoteric decision theory questions, studying trust by looking for decision procedures which are trustworthy in the broadest possible set of decision problems (specifically, attempting to combine Wei Dai's updateless decision theory with learning, and also studying some questions about CDT vs EDT). More recently, I have de-prioritized these questions of perfecting decision theory in favor of trying to model something incrementally closer to what happens in frontier labs. Specifically, this means modeling trust in the updateful case, using Garrabrant Induction (aka Logical Induction) as a bounded-rationality model. One inductor represents human scientific-philosophical progress, while a second one represents AI progress. When should one inductor trust another?
Abram Demski is an AI Safety researcher specializing in Agent Foundations, best known for Embedded Agency (co-written with Scott Garrabrant). His overall approach primarily involves deconfusion research in relation to various concepts related to AI risks, including agency, optimization, trust, meaning, understanding, interpretability, and computational uncertainty (more commonly but less precisely known as bounded rationality). More specifically, his recent work focuses on modeling trust, with the objective of clarifying conditions under which humans can justifiably trust AI.
We can discuss this more and decide on a different structure, but by default, 1 hour 1-on-1 meetings with each scholar once a week, plus a 2 hour group meeting which may also include outside collaborators.
Quality of fit is roughly proportional to philosophical skill times mathematical skill. Someone with excellent philosophical depth and almost no mathematics could be an OK fit, but would probably struggle to produce or evaluate proofs. Someone with excellent mathematical depth but no philosophy could be an OK fit, but might struggle to understand what assumptions and theorems are useful/interesting.
There will be some flexibility about what specific projects scholars will pursue. Abram will discuss the current state of his research with scholars and what topics scholars are interested in, aiming to settle on a topic by or before week 2.
The Winter 2026 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.