MATS mentors are advancing the frontiers of AI alignment, transparency, and security

Lee Sharkey
Goodfire AI
,
Principal Investigator
—

Lee Sharkey is a Principal Investigator at [Goodfire](https://www.goodfire.ai/).

​

His team has focused on improved interpretability methods, including parameter decomposition methods such as [Attribution-based Parameter Decomposition](https://arxiv.org/abs/2501.14926) and [Stochastic Parameter Decomposition](https://arxiv.org/abs/2506.20790) and [adVersarial Parameter Decomposition](https://www.goodfire.ai/research/interpreting-lm-parameters#).

​

Previously, Lee was Chief Strategy Officer and cofounder of [Apollo Research](https://www.apolloresearch.ai/), and a Research Engineer at Conjecture, where he worked on [sparse autoencoders as a solution to representational superposition](https://www.alignmentforum.org/posts/z6QQJbtpkEAX3Aojj/interim-research-report-taking-features-out-of-superposition).

Focus:
实证研究
Technical AI Governance, Interpretability
Cody Rushing
Redwood Research
,
Member of Technical Staff
—

Cody is a member of technical staff at Redwood Research working on AI security.

Focus:
实证研究
AI Control and Monitoring, Forecasting and Strategy, Interpretability
Micah Carroll
OpenAI
,
Member of Technical Staff, Safety Systems
—

Micah is a researcher on OpenAI’s safety team interested in AI deception, scalable oversight, and monitorability. He is on leave from a UC Berkeley PhD focused on AI alignment with influenceable humans, AI manipulation from RL training, and recommender-system effects.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Misalignment Science, Alignment Training Methods, Structural Risk and Societal Dynamics
Robert Kirk
UK AISI
,
Research Scientist
—

Robert is a research scientist and the acting lead of the alignment red-teaming sub-team at UK AISI. This team's focus is on stress-testing model alignment to detect and understand model propensities relevant to loss-of-control risks. Before that, he's most recently worked on misuse research, focusing on evaluations of safeguards against misuse and mitigations for misuse risk, particularly in open-weight systems. He graduated from his PhD from University College London on generalisation in LLM fine-tuning and RL agents in January 2025.

Focus:
实证研究
Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, Misalignment Science

Alex Souly is a researcher on the Red Team at the UK AI Security Institute, where she works on the safety and security of frontier LLMs. She has contributed to pre-deployment evaluations and red-teaming of misuse safeguards and alignment (see Anthropic and OpenAI blogpost), and worked on open source evals like StrongReject and AgentHarm. Previously, she studied Maths at Cambridge and Machine Learning at UCL as part of UCL Dark lab, interned at CHAI, and in another life worked as a SWE at Microsoft.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, Misalignment Science, AI Systems Security
Eric Winsor
UK AISI
,
Research Engineer
—

Eric Winsor is a research scientist at the UK AI Security Institute and contributes to adversarial testing of frontier AI model safeguards. Winsor earned a B.S.E. in computer engineering from the University of Michigan.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Adversarial Robustness and Safeguards

Romeo 正与 AI 未来项目合作,预测详细的 AI 发展情景并制定政策建议。他主要关注算力与安全领域的预测。此前,他曾是 IAPS 政策研究员,并在哈佛大学同时攻读计算机科学硕士学位,研究方向为系统与硬件。

Focus:
战略与预测
Policy and Governance, Technical AI Governance, Forecasting and Strategy, Structural Risk and Societal Dynamics, AI Systems Security
Patrick Butlin
Eleos
,
Senior Research Lead
—

I am a philosopher of mind and a researcher at Eleos AI, where I work on AI consciousness, agency and welfare. Before joining Eleos, I worked at the Future of Humanity Institute and Global Priorities Institute in Oxford. I'm interested in projects including purely philosophical work on the grounds of moral status; research drawing on cognitive science to gain a mechanistic understanding of sentience and agency; and empirical studies that can shed light on welfare-relevant features in AI.

Focus:
Theory
AI Welfare
Matthew Gentzel
Longview Philanthropy
,
Nuclear Weapons Policy Program Officer
—

Matthew Gentzel is a Nuclear Weapons Policy Program Officer at Longview Philanthropy where he works on grantmaking and priorities research related to reducing the risk of large-scale nuclear war. Roughly half of his grantmaking budget concentrates on AI and emerging tech-related nuclear risk issues, where he investigates risks and opportunities related to AI-enabled targeting, information manipulation, and how perceptions of future AI capability impact escalation control in the near-term.

​

His prior work spanned emerging technology threat and policy assessment, with a particular focus on how advancements in AI may shape the future of influence operations, nuclear strategy, and cyber attacks. He has worked as a policy researcher with OpenAI, as an analyst in the US Department of Defense’s Innovation Steering Group, and as a director of research and analysis at the US National Security Commission on Artificial Intelligence.

​

Mr. Gentzel holds an MA in strategic studies and international economics from Johns Hopkins School of Advanced International Studies, and a BS in fire protection engineering from the University of Maryland College Park.

Focus:
政策与治理
Policy and Governance, Forecasting and Strategy, Structural Risk and Societal Dynamics

Wilson Wu 是对齐研究中心(Alignment Research Center,ARC)的研究员。该中心致力于以系统化、具有理论基础的方法开展机械可解释性研究。他此前还研究过其他可解释性方法,包括紧凑证明以及奇异学习理论的应用。

Focus:
Theory
Interpretability, Theoretical Alignment and Formal Methods

常见问题解答

什么是 MATS 项目?
MATS 导师是谁?
MATS 项目的关键日期有哪些?
谁有资格申请?
申请和导师选择流程是怎样的?