Lee Sharkey is a Principal Investigator at [Goodfire](https://www.goodfire.ai/).
His team has focused on improved interpretability methods, including parameter decomposition methods such as [Attribution-based Parameter Decomposition](https://arxiv.org/abs/2501.14926) and [Stochastic Parameter Decomposition](https://arxiv.org/abs/2506.20790) and [adVersarial Parameter Decomposition](https://www.goodfire.ai/research/interpreting-lm-parameters#).
Previously, Lee was Chief Strategy Officer and cofounder of [Apollo Research](https://www.apolloresearch.ai/), and a Research Engineer at Conjecture, where he worked on [sparse autoencoders as a solution to representational superposition](https://www.alignmentforum.org/posts/z6QQJbtpkEAX3Aojj/interim-research-report-taking-features-out-of-superposition).
Cody is a member of technical staff at Redwood Research working on AI security.
Micah is a researcher on OpenAI’s safety team interested in AI deception, scalable oversight, and monitorability. He is on leave from a UC Berkeley PhD focused on AI alignment with influenceable humans, AI manipulation from RL training, and recommender-system effects.
Robert is a research scientist and the acting lead of the alignment red-teaming sub-team at UK AISI. This team's focus is on stress-testing model alignment to detect and understand model propensities relevant to loss-of-control risks. Before that, he's most recently worked on misuse research, focusing on evaluations of safeguards against misuse and mitigations for misuse risk, particularly in open-weight systems. He graduated from his PhD from University College London on generalisation in LLM fine-tuning and RL agents in January 2025.
Alex Souly is a researcher on the Red Team at the UK AI Security Institute, where she works on the safety and security of frontier LLMs. She has contributed to pre-deployment evaluations and red-teaming of misuse safeguards and alignment (see Anthropic and OpenAI blogpost), and worked on open source evals like StrongReject and AgentHarm. Previously, she studied Maths at Cambridge and Machine Learning at UCL as part of UCL Dark lab, interned at CHAI, and in another life worked as a SWE at Microsoft.
Eric Winsor is a research scientist at the UK AI Security Institute and contributes to adversarial testing of frontier AI model safeguards. Winsor earned a B.S.E. in computer engineering from the University of Michigan.
Romeo 正与 AI 未来项目合作,预测详细的 AI 发展情景并制定政策建议。他主要关注算力与安全领域的预测。此前,他曾是 IAPS 政策研究员,并在哈佛大学同时攻读计算机科学硕士学位,研究方向为系统与硬件。
I am a philosopher of mind and a researcher at Eleos AI, where I work on AI consciousness, agency and welfare. Before joining Eleos, I worked at the Future of Humanity Institute and Global Priorities Institute in Oxford. I'm interested in projects including purely philosophical work on the grounds of moral status; research drawing on cognitive science to gain a mechanistic understanding of sentience and agency; and empirical studies that can shed light on welfare-relevant features in AI.
Matthew Gentzel is a Nuclear Weapons Policy Program Officer at Longview Philanthropy where he works on grantmaking and priorities research related to reducing the risk of large-scale nuclear war. Roughly half of his grantmaking budget concentrates on AI and emerging tech-related nuclear risk issues, where he investigates risks and opportunities related to AI-enabled targeting, information manipulation, and how perceptions of future AI capability impact escalation control in the near-term.
His prior work spanned emerging technology threat and policy assessment, with a particular focus on how advancements in AI may shape the future of influence operations, nuclear strategy, and cyber attacks. He has worked as a policy researcher with OpenAI, as an analyst in the US Department of Defense’s Innovation Steering Group, and as a director of research and analysis at the US National Security Commission on Artificial Intelligence.
Mr. Gentzel holds an MA in strategic studies and international economics from Johns Hopkins School of Advanced International Studies, and a BS in fire protection engineering from the University of Maryland College Park.
Wilson Wu 是对齐研究中心(Alignment Research Center,ARC)的研究员。该中心致力于以系统化、具有理论基础的方法开展机械可解释性研究。他此前还研究过其他可解释性方法,包括紧凑证明以及奇异学习理论的应用。
MATS 项目是一项为期 10 周的研究奖学金计划,旨在培养和支持从事人工智能对齐、透明度和安全领域工作的新兴研究人员。研究员将与世界一流的导师合作,获得专门的研究管理支持,并加入位于伯克利、致力于推动人工智能安全与可靠发展的活跃社区。该项目提供开展高影响力研究并开启人工智能安全领域长期职业生涯所需的架构、资源和指导。
MATS 导师均为来自人工智能安全、对齐、治理、领域建设及安全等广泛领域的顶尖研究人员。他们包括学术界人士、行业研究员以及独立专家,负责指导学者开展研究项目、提供反馈,并助力每位学者的研究成长。导师们的专业领域涵盖:
查看 往届及现任导师
关键日期
申请:
主项目将于 9 月 28 日至 12 月 4 日进行,获选研究员的延展阶段将于 12 月开始。
MATS 欢迎来自不同学术和专业背景的申请者——从机器学习、数学和计算机科学,到政策、经济学、物理学、认知科学、生物学和公共卫生,同时也欢迎没有传统研究背景的创业者、运营人员和领域建设者。主要要求是具备为人工智能安全做出贡献的强烈动机,并展现出技术能力、研究潜力或相关的运营经验。具备人工智能安全相关经验会有所帮助,但并非必要条件。