Richard Ngo

My MATS fellows will do philosophical thinking about multi-agent intelligence and how agents change their values. This will likely involve trying to explore and synthesize ideas from game theory, signaling theory, reinforcement learning, and other related domains.

Stream overview

I'd like theoretically-inclined scholars to work with me towards building up a theory of coalitional agency. This work will be highly abstract, philosophical and exploratory; it's mainly suited for people with strong mathematical backgrounds who've read some of my existing writings on this topic and find them interesting.

Mentors

Richard Ngo
Independent
,
Researcher
SF Bay Area
Agent Foundations
Forecasting and Strategy
Structural Risk and Societal Dynamics

I previously worked on the alignment team at DeepMind, and on the governance team at OpenAI. I'm currently an independent researcher focusing on multi-agent intelligence. My research is in the tradition of natural philosophy; I'm trying to develop vague intuitive concepts (like trust, identity, and integrity) to the point where they can serve as seeds for new scientific paradigms.

Read more

Mentorship style

I'll come meet scholars in person around 2 days a week on average. On those days I'll be broadly available for discussions and brainstorming. On other days scholars can message me for guidance (though I'd prefer to spend most of my effort on this during the in-person days).

Fellows we are looking for

My main criterion for selecting scholars will be clarity of reasoning.

You will probably will work with other scholars in the stream.

Project selection

I will talk through project ideas with the scholar.

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting
Oxford
Theory
AI Welfare
SF Bay Area
Control, Model Organisms, Scheming & Deception, Strategy & Forecasting
SF Bay Area
Theory
Interpretability
Tübingen
Dangerous Capability Evals, Agent Foundations, Adversarial Robustness, Monitoring, Scalable Oversight, Scheming & Deception
SF Bay Area
Policy and Governance
Dangerous Capability Evals, Policy & Governance
New York City
Monitoring, Dangerous Capability Evals, Scalable Oversight, Safeguards
SF Bay Area
Policy and Governance
Strategy & Forecasting, Policy & Governance
Montreal
Agent Foundations, Dangerous Capability Evals, Monitoring, Control, Red-Teaming, Scalable Oversight
SF Bay Area
Control, Model Organisms, Red-Teaming, Scheming & Deception