Civilizational Alignment

We aim to generalize tools for analyzing the dynamics of large-scale agency and power, such as public choice theory, to the setting in which machine minds are competitive with humans.

Stream overview

We're open to fellows suggesting projects, but here are some concrete ones that we'd support:

  1. Characterizing AI pacing dynamics. What levers are available to different actors, and what are their effects?
  2. Characterizing the space of possible good futures. Which tradeoffs are inevitable, and where could we look for huge Pareto improvements? E.g. Could sophisticated paternalism allow local control while avoiding disastrous outcomes? Under what conditions should we expect moral convergence?
  3. Fleshing out what "political" natural selection will lead to under various UBI/governance schemes. What mechanisms might states use to tie representation to particular minds, and how might those be predictably hacked? I.e. how will the definition of "person" likely evolve, and then be spammed?

Mentors

Raymond Douglas
ACS Research, University of Toronto
,
Senior Researcher
London
Agent Foundations
Forecasting and Strategy
Structural Risk and Societal Dynamics

My main research interest is figuring out what a good future might look like given the development of very advanced AIs, including how society might be structured and what types of AIs might exist. I also do some empirical research on language model psychology. My first real foray into research was MATS 4.0, focused on theories of agency for predictive models.

Read more
David Duvenaud
University of Toronto
,
Associate Professor
Forecasting and Strategy
Capability and Propensity Evaluations
Structural Risk and Societal Dynamics

David Duvenaud is an Associate Professor in Computer Science and Statistics at the University of Toronto, who now works mainly on problems related to civilizational alignment, i.e. understanding what it will take to keep states and institutions aligned to human interests post-AGI. He holds a Sloan Research Fellowship, a Canada Research Chair in Generative Models, and a CIFAR AI chair. His postdoc was done at Harvard University and his Ph.D. at the University of Cambridge. He is a Founding Member of the Vector Institute for Artificial Intelligence. In 2023-2024 he did a sabbatical at Anthropic, leading their Alignment Evaluations team, as well as research projects on jailbreaks and sabotage. He's also a co-chair of the Schwartz Reisman Institute for Technology and Society, a director of the AI Safety Foundation, and an advisor to AVERI. He has also received a Google Faculty Award, and best paper awards at both the Neural Information Processing Systems (NeurIPS) conference and the International Conference on Machine Learning (ICML).

Read more

Mentorship style

Fellows we are looking for

We're open to all backgrounds. Our ideal candidate might look something like Robin Hanson or David Friedman - a polymath who is comfortable both with analytical tools (e.g. from economics) and with extensive knowledge of real human history, institutions, and the pressures under which populations, cultures, states, and organizations of all sorts evolve.

Project selection

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

Empirical
London
Empirical
London
Empirical
Montreal
Empirical
SF Bay Area
Empirical
Theory
Founding and Field-Building
Policy and Governance
SF Bay Area
Empirical
Systems Security
London
Empirical
Washington, D.C.
Policy and Governance
Washington, D.C.
Policy and Governance
SF Bay Area
Strategy and Forecasting
Policy and Governance
SF Bay Area
Founding and Field-Building
Washington, D.C.
Biosecurity
Empirical
London
Empirical
London
Empirical