MATS mentors are advancing the frontiers of AI alignment, transparency, and security

Alek Westover
Redwood Research
,
Member of Technical Staff
—

Alek is working on AI safety at Redwood Research.  He recently graduated from MIT where he studied Math, CS and AI.  Before working on AI safety he did theoretical computer science research (data structures, online algorithms, and algorithmic graph theory).

Focus:
实证研究
AI Control and Monitoring, Misalignment Science, Alignment Training Methods, Forecasting and Strategy

Ryan is Chief scientist at Redwood Research, focused on technical AI safety research to reduce risks from rogue AIs.

Focus:
实证研究
AI Control and Monitoring, Misalignment Science, Forecasting and Strategy

I am an Assistant Professor of Statistics and EECS at UC Berkeley, where I’m also part of BAIR and CLIMB. I am also Founder & CEO of Transluce, a non-profit research lab building open, scalable technology for understanding frontier AI systems.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, Misalignment Science, Forecasting and Strategy, Interpretability

I'm a research scientist at the UK AI Security Institute, working on AI control red teaming and model organisms of misalignment. I was previously a postdoc with Sam Bowman at NYU, did MATS with Owain Evans, and mentored for the MATS, SPAR and Pivotal fellowships. I got my PhD at the University of Edinburgh, supervised by Iain Murray.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, Misalignment Science
Adam Kaufman
Redwood Research
,
Member of Technical Staff
—

Adam is an AI Safety researcher and member of technical staff at Redwood Research.

Focus:
实证研究
AI Control and Monitoring
Stephen McAleer
Anthropic
,
Member of Technical Staff
—

Stephen is currently a researcher at Anthropic where he researches how to align and control superintelligence. He was previously a researcher at OpenAI and, before that, a postdoc at CMU working with Tuomas Sandholm. Stephen received his PhD in computer science from the University of California, Irvine working with Pierre Baldi. During his PhD, he did research scientist internships at Intel Labs and DeepMind. Before that, Stephen received his bachelor's degree in mathematics and economics from Arizona State University in 2017. Projects he is interested in include:

  • Anything related to control/monitoring for coding agents
  • Scalable oversight for agent alignment
  • Scheming evaluations and mitigations
  • Adversarial training for robust monitors / reward models
  • Reward hacking / deception in agents
Focus:
实证研究
AI Control and Monitoring, Adversarial Robustness and Safeguards, Misalignment Science
Gabriel Kulp
RAND
,
Executive Director
—

Gabriel works with RAND on hands-on projects to build and test prototypes of secure compute infrastructure. He focuses on how to secure the most sensitive AI data centers against the most sophisticated current and future threats. Gabriel has also worked on hardware-enabled governance mechanisms (HEMs, at the intersection of GPU export control and hardware security) and on technical verification of agreements on the development and use of AI systems. He holds a master's degree in computer science and is pursuing a PhD in AI.

Focus:
系统安全
Technical AI Governance, AI Systems Security
Kyle Fish
Anthropic
,
Model Welfare Lead
—

Kyle works on model welfare at Anthropic. He previously co-founded Eleos AI Research, Telis Bioscience, and Alvea.

Focus:
实证研究
Misalignment Science, AI Welfare
Julian Stastny
Redwood Research
,
associate member of technical staff
—

Julian leads the diffuse control team at Redwood.

Focus:
实证研究
AI Control and Monitoring, Misalignment Science, Forecasting and Strategy
Alex Cloud
Anthropic
,
Member of Technical Staff (Alignment Science)
—

Alex is a researcher at Anthropic. He is interested in developing principled methods to induce safety-relevant structure in models. Examples include gradient routing to localize learning updates in models and distillation for robust unlearning.

​

Previously, Alex conducted applied research in reinforcement learning at Riot Games AI and Amazon. He earned a PhD in Statistics from North Carolina State University, where he was advised by Eric Laber.

Focus:
实证研究
Misalignment Science, Alignment Training Methods, Interpretability

常见问题解答

什么是 MATS 项目?
MATS 导师是谁?
MATS 项目的关键日期有哪些?
谁有资格申请?
申请和导师选择流程是怎样的?