Raymond Douglas & David Duvenaud

This stream will explore AI slowdown dynamics, desirable futures, moral convergence, and open-ended forecasting.

Stream overview

Raymond Douglas

I'm pretty opportunistic; things I am currently thinking about which are representative:

  • Should there be a field studying AI slowdown dynamics in general? What should it look like?
  • What do good futures look like, and what are the toughest tensions? E.g. how do you handle paternalism? (I am personally sceptical of handover)
  • How does moral convergence work? Is it a viable win condition?

———

David Duvenaud (co-mentor)

  • Fleshing out what "political" natural selection will result in under various UBI schemes.  I.e. how will people and machines Goodhart various attempts to redistribute wealth.  Baby factories?  Gerrymandering distributed compute to meet the minimal official threshold for consciousness?
  • A more conceptual project outlining the factors pushing for agency at the global level, versus for uncoordinated competition.  I.e. is the usual outcome for a civilization increasing centralization, or a runaway race to the bottom?
  • Working on the technical problem of open-ended forecasting with LLMs, e.g. predicting future newspaper headlines.  A relevant sub-problem is evaluating these abilities on historical data without being confounded by things like shifts in style.

Mentors

Raymond Douglas
ACS Research
,
Senior Researcher
London
Agent Foundations
Forecasting and Strategy
Structural Risk and Societal Dynamics

Raymond Douglas is a senior researcher at ACS Research, studying the societal effects of powerful AI systems and how agency works in AI. During the COVID-19 pandemic, Douglas published epidemiology research and advised the UK government.

Read more
David Duvenaud
University of Toronto
,
Associate Professor
Forecasting and Strategy
Capability and Propensity Evaluations
Structural Risk and Societal Dynamics

David Duvenaud is an associate professor at the University of Toronto, researching AGI governance, evaluation, and catastrophic-risk mitigation. He previously led Anthropic’s Alignment Evaluations team.

Read more

Mentorship style

Fellows we are looking for

Project selection

Streams

The Winter 2026 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

SF Bay Area
Dangerous Capability Evals
Boston
Policy and Governance
Adversarial Robustness, Policy & Governance, Red-Teaming, Safeguards
New York City
Control, Scalable Oversight, Red-Teaming, Model Organisms, Monitoring
SF Bay Area
Policy and Governance
Policy & Governance
SF Bay Area
Control, Monitoring, Dangerous Capability Evals
SF Bay Area
Security, Compute Infrastructure
London
Theory
Interpretability
London
Scheming & Deception, Dangerous Capability Evals, Control, Red-Teaming
SF Bay Area
Dangerous Capability Evals, Red-Teaming, Model Organisms, Control, Monitoring
Toronto
Interpretability
London
Control, Monitoring, Safeguards, Dangerous Capability Evals, Scheming & Deception
Chicago
Biorisk, Security, Safeguards
SF Bay Area
Interpretability, Agent Foundations
London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting