MATS mentors are advancing the frontiers of AI alignment, transparency, and security

Nicholas is a research scientist at Anthropic (previously Google DeepMind) researching adversarial machine learning; he likes to break things.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, AI Systems Security
Sam Bowman
Anthropic
,
Member of Technical Staff
—

Sam Bowman leads a research group working on AI alignment and welfare at Anthropic, with a particular focus on evaluation. Sam is also on leave from NYU as an Associate Prof. of Computer Science and Data Science. He has been studying neural network language models since 2012.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Misalignment Science, AI Welfare
Joe Benton
Anthropic
,
Member of Technical Staff
—

Joe is a member of the Alignment Science team at Anthropic. He's currently working on scalable oversight and also has interests in control, chain-of-thought monitoring, and alignment evaluations. For some examples of recent projects, including MATS collaborations, see: https://joejbenton.com/research/.

Focus:
实证研究
AI Control and Monitoring, Adversarial Robustness and Safeguards, Misalignment Science
Yoshua Bengio
Mila, LawZero
,
Co-President and Scientific Director (LawZero) / Full Professor (UdeM) / Founder and Scientific Advisor (Mila)
—

Yoshua Bengio is Full Professor of Computer Science at Université de Montreal, Co-President and Scientific Director of LawZero, as well as the Founder and Scientific Advisor of Mila. He also holds a Canada CIFAR AI Chair. Considered one of the world’s leaders in Artificial Intelligence and Deep Learning, he is the recipient of the 2018 A.M. Turing Award, considered to be the "Nobel Prize of computing." He is the most cited computer scientist worldwide, and the most-cited living scientist across all fields (by total citations).

​

Professor Bengio is a Fellow of both the Royal Society of London and Canada, an Officer of the Order of Canada, a Knight of the Legion of Honor of France, a member of the UN’s Scientific Advisory Board for Independent Advice on Breakthroughs in Science and Technology, and chairs the International AI Safety Report.

Focus:
实证研究
AI Control and Monitoring, Adversarial Robustness and Safeguards, Policy and Governance, Technical AI Governance, Agent Foundations, Theoretical Alignment and Formal Methods
Alex Turner
Independent
,
Independent researcher
—

Alex is currently working on training invariants into model behavior. In the past, he formulated and proved the power-seeking theorems, co-formulated the shard theory of human value formation, and proposed the Attainable Utility Preservation approach to penalizing negative side effects.

Highlighted outputs from past streams:

​

Focus:
实证研究
Alignment Training Methods, Interpretability, Agent Foundations

Daniel is working on forecasting detailed AI scenarios with Eli Lifland, Thomas Larsen, Jonas Vollmer, and Romeo Dean.

Focus:
战略与预测
Policy and Governance, Technical AI Governance, Forecasting and Strategy, Structural Risk and Societal Dynamics

David Lindner is a Research Scientist on Google DeepMind's AGI Safety and Alignment team where he works on evaluations and mitigations for deceptive alignment and scheming. His recent work includes MONA, a method for reducing multi-turn reward hacking during RL, designing evaluations for stealth and situational awareness, and helping develop GDM's approach to deceptive alignment. Currently, David is interested in studying mitigations for scheming, including CoT monitoring and AI control. You can find more details on his website.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Misalignment Science, Alignment Training Methods

Thomas 是 AI 未来项目的研究员,也是广受关注的 AI 2027 情景预测报告的共同作者。他此前创立了 AI 安全倡导组织 Center for AI Policy,并曾在机器智能研究所开展 AI 安全研究。

Focus:
战略与预测
Policy and Governance, Forecasting and Strategy, Structural Risk and Societal Dynamics
Trenton Bricken
Anthropic
,
Member of Technical Staff
—

I'm a Member of Technical Staff on the Alignment Science team at Anthropic. I'm currently enabling Claude to automatically audit and detect misalignment.

​

About me

  • I have a PhD in Systems Biology from Harvard. My thesis was on "Sparse Representations in Biological and Artificial Neural Networks" in the Kreiman Lab with support from the NSF Graduate Research Fellowship. I also spent time at the Berkeley Redwood Center for Theoretical Neuroscience as a visiting researcher.
  • I graduated from Duke University in May 2020 with a self-made major in "Minds and Machines: Biological and Artificial Intelligence". I was lucky to attend as a Robertson Scholar, which provided full funding during all four years, including summer experiences.
  • At Duke, I spent a year doing research in Dr. Michael Lynch's Lab attempting to use machine learning to design new CRISPR guide RNAs for safer, more effective genome editing. Afterwards, I was affiliated with Dr. Debora Marks's Lab at Harvard Medical School applying deep learning to protein design. I also contributed to the IARPA Fun GCAT and DARPA Biostasis programs.
Focus:
实证研究
AI Control and Monitoring, Misalignment Science, Interpretability
Tyler Tracy
Redwood Research
,
Member of Technical Staff
—

Tyler is an AI Safety Researcher and member of technical staff at Redwood Research.

Focus:
实证研究
AI Control and Monitoring

常见问题解答

什么是 MATS 项目?
MATS 导师是谁?
MATS 项目的关键日期有哪些?
谁有资格申请?
申请和导师选择流程是怎样的?