Mauricio Baker, Anjay Friedman

This stream focuses on AI policy, especially technical governance topics. Tentative project options include: technical projects for verifying AI treaties, metascience for AI safety and governance, and proposals for tracking AI-caused job loss. Scholars can also propose their own projects.

Stream overview

Verifying international agreements on AI:

Verification of compliance could be crucial for international agreements on AI, and ultimately for safe and beneficial AI. Building on a recent paper, various follow-up projects are needed to make verification possible. These could include deep-dives into specific verification mechanisms or protocols, developing proof-of-concept implementations (e.g. of using Confidential Computing to verifiably implement a proof-of-training-data protocol), and further operationalizing problems to advance R&D.

Mentors

Mauricio Baker
RAND, University of Oxford
,
Technical AI Policy Research Scientist; DPhil (PhD) student
Washington, D.C.
Policy and Governance
AI Systems Security
Technical AI Governance

Mauricio researches AI policy at RAND and Oxford. His work has focused on verification of international agreements on AI. He’s more broadly interested in technical AI governance. Previously, Mauricio contracted with OpenAI and did a master's in Computer Science at Stanford University.

Read more

Mentorship style

We'll meet once or twice a week (~1 hr/wk total, as a team if it's a team project). I'm based in DC, so we'll meet remotely. I (Mauricio) will also be available for async discussion, career advising, and detailed feedback on research plans and drafts.

Fellows we are looking for

No hard requirements. Bonus points for research experience, AI safety and governance knowledge, writing and analytical reasoning skills, and experience relevant to specific projects.

Project selection

I'll talk through project ideas with scholar

Streams

The Winter 2026 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

SF Bay Area
Dangerous Capability Evals
Boston
Policy and Governance
Adversarial Robustness, Policy & Governance, Red-Teaming, Safeguards
New York City
Control, Scalable Oversight, Red-Teaming, Model Organisms, Monitoring
SF Bay Area
Policy and Governance
Policy & Governance
SF Bay Area
Control, Monitoring, Dangerous Capability Evals
SF Bay Area
Security, Compute Infrastructure
London
Theory
Interpretability
London
Scheming & Deception, Dangerous Capability Evals, Control, Red-Teaming
SF Bay Area
Dangerous Capability Evals, Red-Teaming, Model Organisms, Control, Monitoring
Toronto
Interpretability
London
Control, Monitoring, Safeguards, Dangerous Capability Evals, Scheming & Deception
Chicago
Biorisk, Security, Safeguards
SF Bay Area
Interpretability, Agent Foundations
London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting