Abram Demski

Agent Foundations research focused on clarifying conditions under which humans can justifiably trust artificial intelligence systems. When should one boundedly rational learning-theoretic process come to trust another?

Stream overview

Methodologically speaking, a proposed AI safety technique is only as good as its safety argument. This does not imply that we need to mathematically prove absolute safety from no assumptions (which is impossible), but it does mean we need to track what assumptions we need to make and rigorously connect them into an argument. The best (most real) risk arguments (eg "If Anyone Builds It, Everyone Dies") and safety arguments (eg basin-of-corrigibility arguments or alignment-by-default via natural abstractions) are still highly informal; I aim to come up with formal arguments, accepting some simplifications at first for the sake of progress, and moving towards increasingly realistic models.

My main line of research tackles this problem head-on by analyzing agentic trust directly, with the aim of clarifying how humans can justifiably trust AI systems.  As of this writing, my current thinking about research directions for this project is described here. An older, more out-of-date but more thorough description of research directions is here. If this line of research goes well, it would produce relevant advice for AI engineering (such as LLM work).

Historically, I have focused on somewhat esoteric decision theory questions, studying trust by looking for decision procedures which are trustworthy in the broadest possible set of decision problems (specifically, attempting to combine Wei Dai's updateless decision theory with learning, and also studying some questions about CDT vs EDT). More recently, I have de-prioritized these questions of perfecting decision theory in favor of trying to model something incrementally closer to what happens in frontier labs. Specifically, this means modeling trust in the updateful case, using Garrabrant Induction (aka Logical Induction) as a bounded-rationality model. One inductor represents human scientific-philosophical progress, while a second one represents AI progress. When should one inductor trust another?

Mentors

Abram Demski
AFFINE
,
Research Scientist
Grand Rapids
Agent Foundations
Theoretical Alignment and Formal Methods

Abram Demski is an AI Safety researcher specializing in Agent Foundations, best known for Embedded Agency (co-written with Scott Garrabrant). His overall approach primarily involves deconfusion research in relation to various concepts related to AI risks, including agency, optimization, trust, meaning, understanding, interpretability, and computational uncertainty (more commonly but less precisely known as bounded rationality). More specifically, his recent work focuses on modeling trust, with the objective of clarifying conditions under which humans can justifiably trust AI.

Read more

Mentorship style

We can discuss this more and decide on a different structure, but by default, 1 hour 1-on-1 meetings with each scholar once a week, plus a 2 hour group meeting which may also include outside collaborators.

Fellows we are looking for

  • Ability to read and write mathematical proofs.
  • Some fluency with probability theory and expected utility theory.

Quality of fit is roughly proportional to philosophical skill times mathematical skill. Someone with excellent philosophical depth and almost no mathematics could be an OK fit, but would probably struggle to produce or evaluate proofs. Someone with excellent mathematical depth but no philosophy could be an OK fit, but might struggle to understand what assumptions and theorems are useful/interesting.

Project selection

There will be some flexibility about what specific projects scholars will pursue. Abram will discuss the current state of his research with scholars and what topics scholars are interested in, aiming to settle on a topic by or before week 2.

Streams

The Winter 2026 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

Systems Security
Systems Security
SF Bay Area
Empirical
London
Empirical
SF Bay Area
Empirical
SF Bay Area
Empirical
SF Bay Area
Founding and Field-Building
Systems Security
Washington, D.C.
Policy and Governance
SF Bay Area
Founding and Field-Building
Biosecurity
London
Theory
London
Empirical
SF Bay Area
Empirical
Theory
SF Bay Area
Strategy and Forecasting
Policy and Governance
No items found.
SF Bay Area
Founding and Field-Building
London
Biosecurity
Washington, D.C.
Biosecurity
Empirical
London
Empirical
London
Empirical