MATS mentors are advancing the frontiers of AI alignment, transparency, and security

Scott Emmons
Anthropic
,
Research Scientist
—

I research AI safety and alignment at Anthropic. Before that, I was a research scientist at Google DeepMind. I completed my PhD at UC Berkeley's Center for Human-Compatible AI, advised by Stuart Russell. I previously cofounded FAR.AI, a 501(c)3 research nonprofit that incubates and accelerates beneficial AI research agendas.

​

I develop AI alignment frameworks, stress-test their limits, and turn insights into methodology adopted across the field. I have established that chain-of-thought monitoring is a substantial defense when reasoning is necessary for misalignment, designed practical metrics to preserve monitorability during model development, shown that obfuscated activations can bypass latent-space defenses, and developed StrongREJECT, a jailbreak benchmark now used by OpenAI, US/UK AISI, Amazon, and others.

Focus:
实证研究
AI Control and Monitoring, Adversarial Robustness and Safeguards, Misalignment Science, Agent Foundations
Krishnamurthy Dvijotham (Dj)
Google DeepMind
,
Senior Staff Research Scientist
—

Krishnamurthy (Dj) Dvijotham is a senior staff research scientist at Google DeepMind, where he leads efforts on the development of secure and trustworthy AI agents. He previously founded the AI security research team at ServiceNow Research and co-founded the robust and verified AI team at DeepMind. His past research has received best paper awards at many leading AI conferences, including most recently at ICML and CVPR 2024. His research led to the framework used for AI security testing at ServiceNow and has been deployed in several Google products, including the Android Play Store, YouTube and Gemini.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, AI Systems Security, Theoretical Alignment and Formal Methods
Dan Mossing
Anthropic
,
Member of technical staff
—

I am an interpretability researcher at Anthropic. I am most interested in simple, practical interpretability approaches that are targeted at making models safer. In a previous life, I worked as a neuroscientist.

Focus:
实证研究
Misalignment Science, Interpretability
Jacob Merizian
UK AISI
,
Research Scientist, Workstream Lead
—

I work at the [UK AI Security Institute](https://www.aisi.gov.uk/). In the past, I’ve done research in high-performance computing, language model pretraining, interpretability, and [hardware enabled governance](https://flexheg.com/).

Focus:
实证研究
AI Control and Monitoring, Misalignment Science, Technical AI Governance
Seth Donoughe
RAND
,
Director of AI
—

Seth is a senior advisor on research and policy at SecureBio, where he works on AI and biosecurity. He previously completed a Ph.D. at Harvard and postdoctoral research at the University of Chicago.

Focus:
实证研究
Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, Biosecurity

Shi Feng leads a research group working on oversight and control. He is an assistant professor at George Washington University. Prior to that, he was a postdoc in the NYU Alignment Research Group under Sam Bowman. He currently focuses on deception and collusion, with an emphasis on propensity and evaluation realism.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Misalignment Science

Adam Shai has extensive research experience in experimental and computational neuroscience. He earned his PhD from Caltech and has over a decade of experience investigating the neural basis of intelligent behavior, most recently as a researcher at Stanford. Driven by the pressing need for AI safety, he has now turned his expertise to neural networks, aiming to develop principled methods for controlling and aligning increasingly advanced AI systems.

​

Adam co-founded and now leads research at Simplex, an organization dedicated to building a science of representations in AI systems.

Focus:
实证研究
Interpretability
Daniel Murfet (Dan)
Timaeus
,
Director of Research
—

I was until recently a professional mathematician at the University of Melbourne, where I worked on algebraic geometry, mathematical logic, some aspects of mathematical physics, and most recently statistical learning theory. As of early 2025 I left academia to direct research at Timaeus on AI safety.

Focus:
实证研究
Alignment Training Methods, Interpretability, Theoretical Alignment and Formal Methods
Saad Siddiqui
Safe AI Forum
,
Senior AI Policy Researcher
—

Saad Siddiqui is a senior researcher at Safe AI Forum, where his research examines possible agreement between leading AI powers. He previously worked as a management consultant at Bain and Company in Singapore.

Focus:
政策与治理
Policy and Governance, Technical AI Governance

Oly works at the Future of Life Foundation on sourcing and developing ambitious ideas to build a flourishing future, grounded in realistic scenarios for AI and other technological development. Priorities include human collective intelligence uplift, gentle and manageable multiagent transitions, and defensive tech.

​

Oly previously worked on loss of control risk modelling and evaluation at the UK AI Safety/Security Institute and continues to engage with the OECD, UK FCDO, DSIT, and parliamentarians on AI governance.

​

He researched (LM) agent oversight and multiagent safety at Oxford and was one of the first beneficiaries of the MATS program in 2021-22. Before his AI safety work, he was a senior data scientist and software engineer.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Policy and Governance, Technical AI Governance, Forecasting and Strategy, Structural Risk and Societal Dynamics, Agent Foundations

常见问题解答

什么是 MATS 项目?
MATS 导师是谁?
MATS 项目的关键日期有哪些?
谁有资格申请?
申请和导师选择流程是怎样的?