Safe AI Forum

International coordination to reduce frontier AI risks, with a focus on China and the West.

Stream overview

There are several clusters of work that scholars could work with us on.

  1. US-China interstate coordination on frontier AI safety. Projects in this vein could include  (a) researching and developing specific proposals for international agreements that could reduce frontier risks, (b) research that advances understanding of how advanced AI and AGI will affect geopolitics, (c) research that increases the likelihood of successful coordination (e.g., verification of AI agreements)
  2. Collaborative frontier AI safety and governance research projects with international experts. These include topics such as frontier AI regulation and loss of control threat modelling. This would likely involve joining an ongoing project and contributing to it rather than starting an independent project.
  3. Horizon scanning and identifying other possible topics for coordination between China and the West.

Mentors

Fynn Heide
Safe AI Forum
,
Executive Director
SF Bay Area
Policy and Governance
Technical AI Governance

Fynn Heide is Executive Director of the Safe AI Forum. Previously, Heide researched AI policy in China as a research scholar at the Centre for the Governance of AI.

Read more
Isabella Duan
Safe AI Forum
,
Senior Researcher
SF Bay Area
Policy and Governance
Technical AI Governance

Isabella is a Senior Researcher at the Safe AI Forum, where she works on U.S.–China coordination on frontier AI safety. Her research focuses on technical governance, particularly building consensus and advancing dialogue on loss-of-control and extreme-misuse risks, with recent work spanning technical misuse safeguards, AI control playbook, and misalignment incidents. She holds an MA in Computational Social Science from the University of Chicago and a BS in Philosophy, Politics, and Economics from University College London.

Read more
Saad Siddiqui
Safe AI Forum
,
Senior AI Policy Researcher
London
Policy and Governance
Technical AI Governance

Saad Siddiqui is a senior researcher at Safe AI Forum, where his research examines possible agreement between leading AI powers. He previously worked as a management consultant at Bain and Company in Singapore.

Read more

Mentorship style

1 hour weekly meetings by default for high-level guidance. We are active on Slack and typically respond within a day for quick questions.

Fellows we are looking for

Good understanding of international AI governance developments that are relevant to frontier AI safety (e.g., the Summit series, AISI network)

Good understanding of Chinese AI governance and safety (key players, key trends and institutional structures)

Good understanding of key frontier risk domains (CBRN, cyber, loss of control)

Some understanding of broader US-China relations and how they frame US-China relations on AI/AGI specifically

Scholars will by default collaborate with someone from the Safe AI Forum, though the exact person will vary based on the project they pick. They are also welcome to work with other collaborators (we are happy to suggest names or have scholars find their own collaborators)

Project selection

We will provide a shortlist of projects that we are keen for the scholar to work on in Week 1. We'll ask scholars to scope these in the 1st week and make a determination about which project to focus on in Week 2.

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting
Oxford
Theory
AI Welfare
SF Bay Area
Control, Model Organisms, Scheming & Deception, Strategy & Forecasting
SF Bay Area
Theory
Interpretability
Tübingen
Dangerous Capability Evals, Agent Foundations, Adversarial Robustness, Monitoring, Scalable Oversight, Scheming & Deception
SF Bay Area
Policy and Governance
Dangerous Capability Evals, Policy & Governance
New York City
Monitoring, Dangerous Capability Evals, Scalable Oversight, Safeguards
SF Bay Area
Policy and Governance
Strategy & Forecasting, Policy & Governance
Montreal
Agent Foundations, Dangerous Capability Evals, Monitoring, Control, Red-Teaming, Scalable Oversight
SF Bay Area
Control, Model Organisms, Red-Teaming, Scheming & Deception