Cristian Trout

In the face of disaster, I suspect the government will be forced to play insurer of last resort, whether for a particular lab, or society at large. (I'm not the only one to suspect this – see e.g. here). Designed well, I believe a federal insurance backstop could internalize catastrophic negative externalities; designed poorly, it will simply be a subsidy for AI companies. I want to design the good version, so we have it ready.

I encourage people with mechanism design (a.k.a. reverse game theory) expertise to apply, but don't be deterred if you don't have this expertise.

Stream overview

I'm trying to design the good version of a US federal insurance backstop for catastrophic risks from AI (read "disasters costing >~$10B"). This draws on precedents like the Price-Anderson Act for nuclear, and the Terrorism Risk Insurance Act.

One version of this research project is quite technical: use reverse game theory to design a mechanism to forecast catastrophic risk from AI. This probably looks like modifying existing forecasting methods like prediction markets and peer prediction mechanisms. The goal: help the government better predict losses, so as to charge risk-priced insurance premiums and incentivize investment in risk mitigations. I believe this is the critical challenge to making a good backstop.

Another version of this project is more like a traditional policy paper. It would step back and compare the different ways the government could play insurer of last resort (e.g. insuring labs vs. reinsuring insurers; pricing the risk in sophisticated ways without requiring additional security measures vs. not attempting risk-pricing but mandating security measures).

Another version of this project is even more high-level, comparing federal insurance backstops with other PPPs that would let the government manage catastrophic risk. I'm much less interested in this version of the project.

Mentors

Cristian Trout
Artificial Intelligence Underwriting Company
,
Research Fellow
SF Bay Area, Washington, D.C.
Policy and Governance
Forecasting and Strategy

Cristian is a Research Fellow at Artificial Intelligence Underwriting Company (AIUC). Insurers have been known to play the role of private regulators (such as in commercial nuclear power); his work broadly focuses on how we might steer the insurance market for AI toward an effective private governance regime.

He was previously a Winter Fellow at the Centre for the Governance of AI, and an independent researcher at the AI Safety Student Team at Harvard. He has an M.A. in Philosophy from the University of British Columbia.

Read more

Mentorship style

1 hour weekly meetings by default for high-level guidance. I'm active on Slack and typically respond within a day for quick questions or conceptual (not code) debugging. Between meetings, expect async back-and-forth on paper structure, or experiment design and results. Scholars can also schedule ad-hoc calls if they're stuck or want to brainstorm—just ping me on Slack.

Depending on the project, I may help with writing.

Fellows we are looking for

If interested in the technical paper, applicants must:

  • Have mechanism design or forecasting design experience (minimum: took graduate level courses on the topic)
  • Have experience with experimental design (e.g. running forecasting surveys)

For all applicants:

Preferred:

  • Published papers in relevant field
  • Graduate degree research experience

Nice to haves:

  • A sense of what makes an actionable, useful policy paper (green flag: you also roll your eyes when you see a paper title like "Towards a taxonomy of frameworks for a principles-based approach to...")
  • Some knowledge of how insurance works
  • Good writing skills

Not a good fit:

  • Scholars who want a lot of freedom over what their research direction will be

Scholars may find their own collaborators if they wish.

Project selection

For technical versions of this project, I suspect the project will automatically be fairly tightly scoped based on the scholar's expertise. I will pose the core challenge and over the first week, the scholar and I will hammer out exactly what theoretical questions need answering + empirical surveys need running.

For non-technical versions of this project, I will pitch a few different projects and scholars will try ones they find interesting for a week. In week 2 we'll settle on one together.

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting
Oxford
Theory
AI Welfare
SF Bay Area
Control, Model Organisms, Scheming & Deception, Strategy & Forecasting
SF Bay Area
Theory
Interpretability
Tübingen
Dangerous Capability Evals, Agent Foundations, Adversarial Robustness, Monitoring, Scalable Oversight, Scheming & Deception
SF Bay Area
Policy and Governance
Dangerous Capability Evals, Policy & Governance
New York City
Monitoring, Dangerous Capability Evals, Scalable Oversight, Safeguards
SF Bay Area
Policy and Governance
Strategy & Forecasting, Policy & Governance
Montreal
Agent Foundations, Dangerous Capability Evals, Monitoring, Control, Red-Teaming, Scalable Oversight
SF Bay Area
Control, Model Organisms, Red-Teaming, Scheming & Deception