本方向涵盖通过机器学习实验开展的实践研究,旨在理解并提升模型安全性,包括 AI 控制、可解释性、可扩展监督、评估、红队测试和鲁棒性。它以研究方法而非单一研究议题为界。如果你主要运用机器学习工程方法,这个方向适合你。
本方向以研究方法而非单一议题为核心。研究员通过机器学习实验,理解并提升前沿模型的安全属性,课题涉及可解释性、AI 控制、可扩展监督、评估、红队测试、鲁棒性,以及错位行为模型生物等。共同特点是通过实际模型开展工作(训练、探测、微调、测量等),而不只是从第一性原理推演。这是项目中规模最大的方向,也是进入技术 AI 安全研究最常见的途径。
我们希望研究员主要使用机器学习工程方法,且具备广义上的相关能力。核心要求是能够设计并运行针对语言模型或其他深度学习系统的实验,并根据结果快速迭代。通常,这意味着熟悉 Python(无论是否借助 AI 编程工具),了解中等规模模型运行所需的基础设施,并能判断哪些实验值得开展。使命契合度也很重要;研究员应能说明某项实证研究如何实质性降低前沿 AI 风险,而不只是它能否产出论文。与其他方向相比,学历和资历并非主要考量。以往的优秀研究员包括本科生,也包括资深行业研究人员。
我们会根据契合度为研究员匹配导师,并规划项目,使其在项目结束前产出具体成果,例如论文、评估套件、开源工具或技术报告。本方向的成果面向前沿实验室的安全与对齐团队、政府及其他评估机构,以及更广泛的机器学习研究社区。
如果你对这类研究感兴趣,欢迎申请。
At Fourth Eon Biosecurity we're building adaptive, AI-native safeguards across the bioengineering stack, with a focus on function-based DNA synthesis screening. Fellows in this stream will work on technical research projects at the intersection of AI safety and biosecurity, aimed at reinforcing screening and generalizing detection beyond known threat signatures. Projects span mechanistic interpretability of bio foundation models, model evaluations for biosecurity-relevant capabilities, and agentic sequence analysis workflows.
I typically schedule a standing weekly 1:1 meeting with each fellow, and also hold a weekly research group meeting. Beyond that I am available on Slack and can find additional time for calls outside of scheduled meetings.
Note that as part of our Safe and Responsible Research Framework we require fellows to sign a fellowship agreement covering confidentiality and pre-publication review for dual-use risks. This is common practice in biosecurity research and allows us to work freely together on sensitive material.
Fellows who are interested in our research area should think of potential project ideas that leverage their strengths and interests. I will work with individual fellows to identify a specific project that matches their background and interests and is aligned with our overall research direction, and to refine the scope and objectives of the project.
I'm interested in better understanding and controlling how post-training causes alignment-relevant behavior. This is a pretty broad area, and I’m open to many approaches to these problems! Potential areas of study / methods of attack might include model organisms, training run science/ablations, root causing strange behaviors, or studying how best to robustly induce behaviors or values or beliefs into models.
The stream focuses on evaluating and/or mitigating catastrophic risk emerging from dangerous scientific capabilities in frontier AI systems, with an emphasis on the challenges that emerge from lab integrations and novel science. Potential research directions include evaluation design, risk mitigations and evaluation science.
We can schedule a weekly 1h meeting, for general progress updates, sharing results, and overall guidance. I would be reachable on Slack as well for async comms. Happy to jump on ad-hoc calls for specific discussions or pair coding/debugging. I am based in London and I work UK hours (10am-7pm), but I also visit the US (Boston) a few times a year.
I will work with the fellow to find the right project that suits their interest within the directions spelled out above. I will pitch a few project ideas and support the fellow in making the decision. I also welcome project suggestions; in those cases I would work with the fellow to scope it appropriately.
This stream focuses on critical challenges in AI safety and alignment, including risks from automating AI research, bottlenecks to recursive self-improvement, and the automation of safety and alignment research. Priority topics also include AGI privacy, measuring long-horizon agentic capabilities, developing new alignment methods, and advancing the science of post-training.
I usually spend at least 30 min per week in one-on-one meetings with my mentees. We can also discuss longer time slots if necessary. Besides these time slots, I try to be as responsive as possible over Slack (>2 comprehensive responses per day) and read relevant papers between weekly meetings.
I would prefer to set the overall direction, but I will listen closely to scholars about their preferences within a broad direction. Converging on a particular topic is expected to be a collaborative process.
Research papers (technical governance or ML) related to evaluating and mitigating dangerous AI capabilities, with a focus on what's actionable and relevant for AGI companies
I like to get daily standup messages about progress that has been made on the project, and I'm happy to provide some quick async feedback on new outputs. I'll also have weekly meetings. I'm based in Constellation in Berkeley.
Good writers/researchers who can work independently and autonomously! I'm looking for scholars who can ship a meaningful research output end-to-end and ideally have prior experience in writing relevant papers.
I may assign a project, have you pick from a list of projects, or talk through project ideas with you.
Neel takes a pragmatic approach to interpretability: identify what stands between where we are now and where we want to be by AGI, and then focus on the subset of resulting research problems that can be tractably studied on today's models. This can look like diving deep into the internals of the model, or simpler black box methods like reading and carefully intervening on the chain of thought - whatever is the right tool for the job. This could look like studying how to detect deception, understanding why a model took a seemingly concerning action, or fixing weak points in other areas of safety, e.g. using interpretability to stop models realising they are being tested. You can learn more about Neel's approach in this podcast.
He has spent far too much time having MATS scholars, and has worked with ~60 so far - he’s excited to take on even more!
Computational/modelling problems in biosecurity.
Typically 1 hour weekly meetings by default. I typically respond on slack quite quickly - some weeks I am not available. You are welcome to chat to my phd students too!
Computational experience e.g. Python OR statistical modelling interest in biosecurity
We will construct a project together that best suits the skills and interests of the fellow and what I can reasonably be helpful for.
Projects in this stream will be on AI welfare and moral status; more specifically, on what it takes to be a moral patient and how we can determine whether AI systems meet the conditions. I'm looking for applicants who have ideas about these topics and are motivated to explore them in more detail.
By default, scholars will meet with me online for 1hr/week and I will respond to questions on email/slack.
I will talk through project ideas with scholar
MATS 项目是一项为期 10 周的研究奖学金计划,旨在培养和支持从事人工智能对齐、透明度和安全领域工作的新兴研究人员。研究员将与世界一流的导师合作,获得专门的研究管理支持,并加入位于伯克利、致力于推动人工智能安全与可靠发展的活跃社区。该项目提供开展高影响力研究并开启人工智能安全领域长期职业生涯所需的架构、资源和指导。
MATS 导师均为来自人工智能安全、对齐、治理、领域建设及安全等广泛领域的顶尖研究人员。他们包括学术界人士、行业研究员以及独立专家,负责指导学者开展研究项目、提供反馈,并助力每位学者的研究成长。导师们的专业领域涵盖:
查看 往届及现任导师
关键日期
申请:
主项目将于 9 月 28 日至 12 月 4 日进行,获选研究员的延展阶段将于 12 月开始。
MATS 欢迎来自不同学术和专业背景的申请者——从机器学习、数学和计算机科学,到政策、经济学、物理学、认知科学、生物学和公共卫生,同时也欢迎没有传统研究背景的创业者、运营人员和领域建设者。主要要求是具备为人工智能安全做出贡献的强烈动机,并展现出技术能力、研究潜力或相关的运营经验。具备人工智能安全相关经验会有所帮助,但并非必要条件。