MATS mentors are advancing the frontiers of AI alignment, transparency, and security

Alex Cloud
Anthropic
,
Member of Technical Staff (Alignment Science)

Alex is a researcher at Anthropic. He is interested in developing principled methods to induce safety-relevant structure in models. Examples include gradient routing to localize learning updates in models and distillation for robust unlearning.

Previously, Alex conducted applied research in reinforcement learning at Riot Games AI and Amazon. He earned a PhD in Statistics from North Carolina State University, where he was advised by Eric Laber.

Focus:
Empirical
Misalignment Science, Alignment Training Methods, Interpretability
Jacob Merizian
UK AISI
,
Research Scientist, Workstream Lead

I work at the UK AI Security Institute. In the past, I’ve done research in high-performance computing, language model pretraining, interpretability, and hardware enabled governance.

Focus:
Empirical
AI Control and Monitoring, Misalignment Science, Technical AI Governance
Lee Sharkey
Goodfire AI
,
Principal Investigator

Lee Sharkey is a Principal Investigator at Goodfire.

His team has focused on improved interpretability methods, including parameter decomposition methods such as Attribution-based Parameter Decomposition and Stochastic Parameter Decomposition and adVersarial Parameter Decomposition.

Previously, Lee was Chief Strategy Officer and cofounder of Apollo Research, and a Research Engineer at Conjecture, where he worked on sparse autoencoders as a solution to representational superposition.

Focus:
Empirical
Technical AI Governance, Interpretability
Cody Rushing
Redwood Research
,
Member of Technical Staff

Cody is a member of technical staff at Redwood Research working on AI security.

Focus:
Empirical
AI Control and Monitoring, Forecasting and Strategy, Interpretability
Micah Carroll
OpenAI
,
Member of Technical Staff, Safety Systems

Micah is a researcher on OpenAI’s safety team interested in AI deception, scalable oversight, and monitorability. He is on leave from a UC Berkeley PhD focused on AI alignment with influenceable humans, AI manipulation from RL training, and recommender-system effects.

Focus:
Empirical
AI Control and Monitoring, Capability and Propensity Evaluations, Misalignment Science, Alignment Training Methods, Structural Risk and Societal Dynamics
Robert Kirk
UK AISI
,
Research Scientist

Robert is a research scientist and the acting lead of the alignment red-teaming sub-team at UK AISI. This team's focus is on stress-testing model alignment to detect and understand model propensities relevant to loss-of-control risks. Before that, he's most recently worked on misuse research, focusing on evaluations of safeguards against misuse and mitigations for misuse risk, particularly in open-weight systems. He graduated from his PhD from University College London on generalisation in LLM fine-tuning and RL agents in January 2025.

Focus:
Empirical
Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, Misalignment Science

Alex Souly is a researcher on the Red Team at the UK AI Security Institute, where she works on the safety and security of frontier LLMs. She has contributed to pre-deployment evaluations and red-teaming of misuse safeguards and alignment (see Anthropic and OpenAI blogpost), and worked on open source evals like StrongReject and AgentHarm. Previously, she studied Maths at Cambridge and Machine Learning at UCL as part of UCL Dark lab, interned at CHAI, and in another life worked as a SWE at Microsoft.

Focus:
Empirical
AI Control and Monitoring, Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, Misalignment Science, AI Systems Security
Eric Winsor
UK AISI
,
Research Engineer

Eric Winsor is a research scientist at the UK AI Security Institute and contributes to adversarial testing of frontier AI model safeguards. Winsor earned a B.S.E. in computer engineering from the University of Michigan.

Focus:
Empirical
AI Control and Monitoring, Capability and Propensity Evaluations, Adversarial Robustness and Safeguards

Romeo is working on forecasting detailed AI scenarios and developing policy recommendations with the AI Futures Project. He focuses primarily on compute and security forecasting. Previously he was an IAPS Policy Fellow and graduated with a concurrent master's in Computer Science at Harvard with a systems and hardware focus.

Focus:
Strategy and Forecasting
Policy and Governance, Technical AI Governance, Forecasting and Strategy, Structural Risk and Societal Dynamics, AI Systems Security
Patrick Butlin
Eleos
,
Senior Research Lead

I am a philosopher of mind and a researcher at Eleos AI, where I work on AI consciousness, agency and welfare. Before joining Eleos, I worked at the Future of Humanity Institute and Global Priorities Institute in Oxford. I'm interested in projects including purely philosophical work on the grounds of moral status; research drawing on cognitive science to gain a mechanistic understanding of sentience and agency; and empirical studies that can shed light on welfare-relevant features in AI.

Focus:
Empirical
AI Welfare

Frequently asked questions

What is the MATS Program?
Who are the MATS Mentors?
What are the key dates of the MATS Program?
Who is eligible to apply?
How does the application and mentor selection process work?