本方向涵盖通过机器学习实验开展的实践研究,旨在理解并提升模型安全性,包括 AI 控制、可解释性、可扩展监督、评估、红队测试和鲁棒性。它以研究方法而非单一研究议题为界。如果你主要运用机器学习工程方法,这个方向适合你。
本方向以研究方法而非单一议题为核心。研究员通过机器学习实验,理解并提升前沿模型的安全属性,课题涉及可解释性、AI 控制、可扩展监督、评估、红队测试、鲁棒性,以及错位行为模型生物等。共同特点是通过实际模型开展工作(训练、探测、微调、测量等),而不只是从第一性原理推演。这是项目中规模最大的方向,也是进入技术 AI 安全研究最常见的途径。
我们希望研究员主要使用机器学习工程方法,且具备广义上的相关能力。核心要求是能够设计并运行针对语言模型或其他深度学习系统的实验,并根据结果快速迭代。通常,这意味着熟悉 Python(无论是否借助 AI 编程工具),了解中等规模模型运行所需的基础设施,并能判断哪些实验值得开展。使命契合度也很重要;研究员应能说明某项实证研究如何实质性降低前沿 AI 风险,而不只是它能否产出论文。与其他方向相比,学历和资历并非主要考量。以往的优秀研究员包括本科生,也包括资深行业研究人员。
我们会根据契合度为研究员匹配导师,并规划项目,使其在项目结束前产出具体成果,例如论文、评估套件、开源工具或技术报告。本方向的成果面向前沿实验室的安全与对齐团队、政府及其他评估机构,以及更广泛的机器学习研究社区。
如果你对这类研究感兴趣,欢迎申请。
The Redwood Research stream is looking for fast empirical iterators and strategists to work on control research.
Depending on the mentor:
We are looking for people who are:
We will assign projects by default but are open to getting pitched on projects.
Roger Grosse’s stream investigates how to improve influence functions and other training data attribution methods, and uses these tools to study alignment-related phenomena such as out-of-context reasoning and emergent misalignment. The ideal scholar has experience with LLM internals, strong statistics/applied math skills (especially numerical linear algebra), and can independently drive research from literature review through experimentation and analysis. Roger provides shovel-ready projects while giving exceptional scholars freedom to pursue their own ideas, and is open to scholars collaborating with others.
I will meet with scholars 1 hour per week by default, and will be available to answer questions on Slack roughly daily.
I will give the scholar the level of freedom they are ready for. I will be prepared with focused, shovel-ready projects, but exceptional scholars with a vision they are excited about will have the flexibility to pursue it.
This stream will work on projects that empirically assess national security threats of AI misuse (CBRN terrorism and cyberattacks) and improve dangerous capability evaluations. Threat modeling applicants should have a skeptical mindset, enjoy case study work, and be strong written communicators. Eval applicants should be able and excited to help demonstrate concepts like sandbagging elicitation gaps in an AI misuse context.
Typically, this would include weekly meetings, detailed comments on drafts, and asynchronous messaging.
For threat modeling work:
For evaluations, mitigations, and verification work:
Mentor(s) will talk through project ideas with scholar
I work on the science of evaluating advanced AI systems for biological and CBRN risks, with a particular interest in translating technical evidence into decisions by governments and frontier AI developers. In this stream, I’m interested in developing novel capability evaluations, studying how dangerous or dual-use capabilities diffuse into increasingly accessible models, and building scalable red-teaming methods that produce rigorous, decision-relevant evidence without requiring risky real-world demonstrations.
In the shard theory stream, we create qualitatively new methods and fields of inquiry, from steering vectors to gradient routing to unsupervised capability elicitation to robust unlearning. If you're theory-minded, maybe you'll help us formalize shard theory itself.
We will have weekly 1-1's and weekly team lunch, as well as asynchronous communication over Slack. Mentees are always welcome to reach out at any time, in case guidance is needed outside of usual meeting times.
Scholars should mostly figure things out on their own outside of meetings
Ideal candidates would have:
Mentor(s) will talk through project ideas with scholar
We are interested in AI control and scalable oversight. I'm excited to work with scholars interested in empirical projects building and evaluating control measures and oversight techniques for LLM agents, especially those based on chain of thought monitoring. I'm also interested in the science of chain of thought monitorability, misalignment and control. An ideal project ends with a paper submitted to NeurIPS/ICML/ICLR.
I'll meet with mentees once a week and will be available on Slack daily.
An ideal mentee has a strong AI research and/or software engineering background. A mentee can be a PhD student and they can work on a paper that will be part of their thesis.
I'll talk through project ideas with scholar
We build scalable technology for AI understanding and oversight.
You will work closely with a mentor through recurring meetings (group and individual) and Slack.
We're looking for strong, experienced software engineers or talented researchers who can hit the ground running and iterate quickly.
We will talk through project ideas with scholars
This stream will focus on evaluating dangerous capabilities in language models and detecting deception and dishonesty.
MATS 项目是一项为期 10 周的研究奖学金计划,旨在培养和支持从事人工智能对齐、透明度和安全领域工作的新兴研究人员。研究员将与世界一流的导师合作,获得专门的研究管理支持,并加入位于伯克利、致力于推动人工智能安全与可靠发展的活跃社区。该项目提供开展高影响力研究并开启人工智能安全领域长期职业生涯所需的架构、资源和指导。
MATS 导师均为来自人工智能安全、对齐、治理、领域建设及安全等广泛领域的顶尖研究人员。他们包括学术界人士、行业研究员以及独立专家,负责指导学者开展研究项目、提供反馈,并助力每位学者的研究成长。导师们的专业领域涵盖:
查看 往届及现任导师
关键日期
申请:
主项目将于 9 月 28 日至 12 月 4 日进行,获选研究员的延展阶段将于 12 月开始。
MATS 欢迎来自不同学术和专业背景的申请者——从机器学习、数学和计算机科学,到政策、经济学、物理学、认知科学、生物学和公共卫生,同时也欢迎没有传统研究背景的创业者、运营人员和领域建设者。主要要求是具备为人工智能安全做出贡献的强烈动机,并展现出技术能力、研究潜力或相关的运营经验。具备人工智能安全相关经验会有所帮助,但并非必要条件。