MATS mentors are advancing the frontiers of AI alignment, transparency, and security

Megan Kinniment
METR
,
Member of Technical Staff

I am a researcher at METR.

I think the development of AI is going to be a confusing time for the world. I want to help provide good evidence and methodologies for tracking AI development and risk, so humanity can make sensible decisions.

I've had different roles at different times, including leading task development and our monitoring stream. I like prototyping new kinds of evaluations. I think it's healthy to read transcripts. I'm interested in what capabilities matter for being a competent agent, and why current AI agents fall short. I feel lucky that I get to spend time building an understanding of the models.

I've previously spent time at the Centre on Long-Term Risk and FHI. Before that I studied physics at university, where I did malaria diagnostics research.

Focus:
Empirical
AI Control and Monitoring, Capability and Propensity Evaluations

Peter is an assistant professor at Princeton University, where he works on reinforcement learning, alignment, and law. He received a J.D. and Ph.D. in computer science from Stanford University.

Focus:
Empirical
Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, Alignment Training Methods, Policy and Governance, Technical AI Governance, Structural Risk and Societal Dynamics

Cristian is a Research Fellow at Artificial Intelligence Underwriting Company (AIUC). Insurers have been known to play the role of private regulators (such as in commercial nuclear power); his work broadly focuses on how we might steer the insurance market for AI toward an effective private governance regime.

He was previously a Winter Fellow at the Centre for the Governance of AI, and an independent researcher at the AI Safety Student Team at Harvard. He has an M.A. in Philosophy from the University of British Columbia.

Focus:
Policy and Governance
Policy and Governance, Forecasting and Strategy
Milad Nasr
Anthropic
,
Research Scientist

Milad is a research scientist at Anthropic studying how language models affect computer security. Before joining Anthropic, Milad researched AI security and privacy at OpenAI and Google DeepMind.

Focus:
Empirical
Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, AI Systems Security
Scott Emmons
Anthropic
,
Research Scientist

I research AI safety and alignment at Anthropic. Before that, I was a research scientist at Google DeepMind. I completed my PhD at UC Berkeley's Center for Human-Compatible AI, advised by Stuart Russell. I previously cofounded FAR.AI, a 501(c)3 research nonprofit that incubates and accelerates beneficial AI research agendas.

I develop AI alignment frameworks, stress-test their limits, and turn insights into methodology adopted across the field. I have established that chain-of-thought monitoring is a substantial defense when reasoning is necessary for misalignment, designed practical metrics to preserve monitorability during model development, shown that obfuscated activations can bypass latent-space defenses, and developed StrongREJECT, a jailbreak benchmark now used by OpenAI, US/UK AISI, Amazon, and others.

Focus:
Empirical
AI Control and Monitoring, Adversarial Robustness and Safeguards, Misalignment Science, Agent Foundations
Krishnamurthy Dvijotham (Dj)
Google DeepMind
,
Senior Staff Research Scientist

Krishnamurthy (Dj) Dvijotham is a senior staff research scientist at Google DeepMind, where he leads efforts on the development of secure and trustworthy AI agents. He previously founded the AI security research team at ServiceNow Research and co-founded the robust and verified AI team at DeepMind. His past research has received best paper awards at many leading AI conferences, including most recently at ICML and CVPR 2024. His research led to the framework used for AI security testing at ServiceNow and has been deployed in several Google products, including the Android Play Store, YouTube and Gemini.

Focus:
Empirical
AI Control and Monitoring, Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, AI Systems Security, Theoretical Alignment and Formal Methods
Dan Mossing
Anthropic
,
Member of technical staff

I am an interpretability researcher at Anthropic. I am most interested in simple, practical interpretability approaches that are targeted at making models safer. In a previous life, I worked as a neuroscientist.

Focus:
Empirical
Misalignment Science, Interpretability
Seth Donoughe
RAND
,
Director of AI

Seth is a senior advisor on research and policy at SecureBio, where he works on AI and biosecurity. He previously completed a Ph.D. at Harvard and postdoctoral research at the University of Chicago.

Focus:
Empirical
Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, Biosecurity
Daniel Murfet (Dan)
Timaeus
,
Director of Research

I was until recently a professional mathematician at the University of Melbourne, where I worked on algebraic geometry, mathematical logic, some aspects of mathematical physics, and most recently statistical learning theory. As of early 2025 I left academia to direct research at Timaeus on AI safety.

Focus:
Empirical
Alignment Training Methods, Interpretability, Theoretical Alignment and Formal Methods
Saad Siddiqui
Safe AI Forum
,
Senior AI Policy Researcher

Saad Siddiqui is a senior researcher at Safe AI Forum, where his research examines possible agreement between leading AI powers. He previously worked as a management consultant at Bain and Company in Singapore.

Focus:
Policy and Governance
Policy and Governance, Technical AI Governance

Frequently asked questions

What is the MATS Program?
Who are the MATS Mentors?
What are the key dates of the MATS Program?
Who is eligible to apply?
How does the application and mentor selection process work?