实证研究

本方向涵盖通过机器学习实验开展的实践研究,旨在理解并提升模型安全性,包括 AI 控制、可解释性、可扩展监督、评估、红队测试和鲁棒性。它以研究方法而非单一研究议题为界。如果你主要运用机器学习工程方法,这个方向适合你。

‍

Application process

  • 第一阶段:完成通用申请
  • 第二阶段第一部分:完成 1 至 2 项评估研究判断力和技术实现能力的测评,并可能需要提供推荐人信息
  • 第二阶段第二部分:回答研究流方向选择问题
  • 第三阶段:参加面试和工作测试

实证研究 track overview

本方向以研究方法而非单一议题为核心。研究员通过机器学习实验,理解并提升前沿模型的安全属性,课题涉及可解释性、AI 控制、可扩展监督、评估、红队测试、鲁棒性,以及错位行为模型生物等。共同特点是通过实际模型开展工作(训练、探测、微调、测量等),而不只是从第一性原理推演。这是项目中规模最大的方向,也是进入技术 AI 安全研究最常见的途径。

我们希望研究员主要使用机器学习工程方法,且具备广义上的相关能力。核心要求是能够设计并运行针对语言模型或其他深度学习系统的实验,并根据结果快速迭代。通常,这意味着熟悉 Python(无论是否借助 AI 编程工具),了解中等规模模型运行所需的基础设施,并能判断哪些实验值得开展。使命契合度也很重要;研究员应能说明某项实证研究如何实质性降低前沿 AI 风险,而不只是它能否产出论文。与其他方向相比,学历和资历并非主要考量。以往的优秀研究员包括本科生,也包括资深行业研究人员。

我们会根据契合度为研究员匹配导师,并规划项目,使其在项目结束前产出具体成果,例如论文、评估套件、开源工具或技术报告。本方向的成果面向前沿实验室的安全与对齐团队、政府及其他评估机构,以及更广泛的机器学习研究社区。

如果你对这类研究感兴趣,欢迎申请。

实证研究 track streams

At Fourth Eon Biosecurity we're building adaptive, AI-native safeguards across the bioengineering stack, with a focus on function-based DNA synthesis screening. Fellows in this stream will work on technical research projects at the intersection of AI safety and biosecurity, aimed at reinforcing screening and generalizing detection beyond known threat signatures. Projects span mechanistic interpretability of bio foundation models, model evaluations for biosecurity-relevant capabilities, and agentic sequence analysis workflows.

Read more
Mentorship structure
Desired fellow characteristics
Project selection process

I'm interested in better understanding and controlling how post-training causes alignment-relevant behavior. This is a pretty broad area, and I’m open to many approaches to these problems! Potential areas of study / methods of attack might include model organisms, training run science/ablations, root causing strange behaviors, or studying how best to robustly induce behaviors or values or beliefs into models.

Read more
Desired fellow characteristics

The stream focuses on evaluating and/or mitigating catastrophic risk emerging from dangerous scientific capabilities in frontier AI systems, with an emphasis on the challenges that emerge from lab integrations and novel science. Potential research directions include evaluation design, risk mitigations and evaluation science.

Read more
Mentorship structure
Desired fellow characteristics
Project selection process

This stream focuses on critical challenges in AI safety and alignment, including risks from automating AI research, bottlenecks to recursive self-improvement, and the automation of safety and alignment research. Priority topics also include AGI privacy, measuring long-horizon agentic capabilities, developing new alignment methods, and advancing the science of post-training.

Read more
Mentorship structure
Desired fellow characteristics
Project selection process

Research papers (technical governance or ML) related to evaluating and mitigating dangerous AI capabilities, with a focus on what's actionable and relevant for AGI companies

Read more
Mentorship structure
Desired fellow characteristics
Project selection process

Neel takes a pragmatic approach to interpretability: identify what stands between where we are now and where we want to be by AGI, and then focus on the subset of resulting research problems that can be tractably studied on today's models. This can look like diving deep into the internals of the model, or simpler black box methods like reading and carefully intervening on the chain of thought - whatever is the right tool for the job. This could look like studying how to detect deception, understanding why a model took a seemingly concerning action, or fixing weak points in other areas of safety, e.g. using interpretability to stop models realising they are being tested. You can learn more about Neel's approach in this podcast.

He has spent far too much time having MATS scholars, and has worked with ~60 so far - he’s excited to take on even more!

Read more

Computational/modelling problems in biosecurity.

Read more
Mentorship structure
Desired fellow characteristics
Project selection process

Projects in this stream will be on AI welfare and moral status; more specifically, on what it takes to be a moral patient and how we can determine whether AI systems meet the conditions. I'm looking for applicants who have ideas about these topics and are motivated to explore them in more detail.

Read more
Mentorship structure
Desired fellow characteristics
Project selection process

常见问题解答

什么是 MATS 项目?
MATS 导师是谁?
MATS 项目的关键日期有哪些?
谁有资格申请?
申请和导师选择流程是怎样的?