本方向涵盖通过机器学习实验开展的实践研究,旨在理解并提升模型安全性,包括 AI 控制、可解释性、可扩展监督、评估、红队测试和鲁棒性。它以研究方法而非单一研究议题为界。如果你主要运用机器学习工程方法,这个方向适合你。
本方向以研究方法而非单一议题为核心。研究员通过机器学习实验,理解并提升前沿模型的安全属性,课题涉及可解释性、AI 控制、可扩展监督、评估、红队测试、鲁棒性,以及错位行为模型生物等。共同特点是通过实际模型开展工作(训练、探测、微调、测量等),而不只是从第一性原理推演。这是项目中规模最大的方向,也是进入技术 AI 安全研究最常见的途径。
我们希望研究员主要使用机器学习工程方法,且具备广义上的相关能力。核心要求是能够设计并运行针对语言模型或其他深度学习系统的实验,并根据结果快速迭代。通常,这意味着熟悉 Python(无论是否借助 AI 编程工具),了解中等规模模型运行所需的基础设施,并能判断哪些实验值得开展。使命契合度也很重要;研究员应能说明某项实证研究如何实质性降低前沿 AI 风险,而不只是它能否产出论文。与其他方向相比,学历和资历并非主要考量。以往的优秀研究员包括本科生,也包括资深行业研究人员。
我们会根据契合度为研究员匹配导师,并规划项目,使其在项目结束前产出具体成果,例如论文、评估套件、开源工具或技术报告。本方向的成果面向前沿实验室的安全与对齐团队、政府及其他评估机构,以及更广泛的机器学习研究社区。
如果你对这类研究感兴趣,欢迎申请。
This stream focuses on secret loyalties, where an LLM covertly tries to advance a principal's interests. Secret loyalties have been established as a pressing threat [1], and model organisms of narrow secret loyalties have been constructed and audited [2]. This stream aims to advance the empirical foundations of our understanding of secret loyalties. The goal being that humanity is well-equipped to deal with secret loyalty installation attempts as and when catastrophic secret loyalties become possible in the future.
We think being well-equipped looks like having sufficient security measures in place in frontier AI companies, understanding the dynamics and behaviours of secretly loyal AI systems, having effective auditing and verification protocols for secret loyalties and attempts to install them, and these protocols actually being followed by relevant stakeholders.
This coalition of mentors make up the “Anthropic Stream”. This stream spans a range of empirical research areas in AI safety on LLMs, including AI control, scalable oversight, model organisms, model internals, model welfare, security, and more. You’ll be pitched, and have the option to pitch, a variety of safety research projects, and then be matched to projects and mentors based on your interests/preferences on research and what you’d like to get out of MATS. Fellows in this stream frequently receive funding and continued mentorship after MATS to complete their research project, usually leading to a (co-)first author paper. People in this stream often end up in long-term homes for safety research after MATS (e.g. Anthropic, Redwood Research, OpenAI).
Anthropic mentors share an application, tend to collaborate and co-mentor projects together, and generally share infrastructure to streamline the fellow experience. By applying to this stream, you are being considered for all of the Anthropic mentors.
During the program, scholars meet weekly with their project mentors and collaborators. Some projects meet more often without mentors (e.g., daily standups with the peers on the project). Each project will have a primary mentor, who is also the main decision-maker on key milestones for the project and who is the default person to go to for feedback, advice, etc. Co-mentors also attend project meetings as needed and provide feedback throughout the program. Some project co-mentors can be as involved as the primary mentor.
Mentorship starts with the “Project Pitch Session” Anthropic runs at the start of the program. Fellows get ~1 week to derisk and trial projects before submitting their preferences. Starting on week 2, scholars are assigned projects where the primary mentor is whoever pitched it. Some projects are assigned co-mentors who are other supervisors who want to join the project.
We will continue working on black-box monitors for scheming in complex agentic settings, building on the success of the previous stream. Concretely, we will work on scaling our datasets and fine-tuning efforts, as described in the scalable monitoring agenda
Most likely the next projects will be about automated iterated red-team vs. blue-team games. We are currently training the blue team. We will then train the red-team and within this stream, we will try and close the loop to train them both synchronously.
We have two weekly 60-minute calls by default. Since everyone will work on the same project, these calls will be with all participants of the stream. I respond on slack on a daily basis for asynchronous messages. Scholars will have a lot of freedom for day-to-day decisions and direction setting. In the best case, you will understand the project better than me after a few weeks and have a clear vision for where it should be heading. I recommend scholars focus 100% of their work time on the project and not pursue anything on the side. I think this way people will learn the most in MATS.
You will work on subprojects of black box monitoring. See here for details.
This stream focuses on building a Science of Scheming, i.e. what are the mechanisms by which future models might become schemers, even though current models are not. We want to discover empirical Scaling Trends for Scheming. For example, does deceptive alignment become easier to discover with improved model capabilities?
1 hour weekly meetings by default for high-level guidance. We’re active on Slack and typically respond within a day for questions. Expect async back-and-forth on experiment design and results between meetings. Scholars can also schedule ad-hoc calls if they're stuck or want to brainstorm—just ping on Slack.
We will set the high-level project direction, as described above. It's not fully clear what exactly the project will look like by the time you start in September. All projects will be in the direction of the Science of Scheming post.
You’d work with the two of us, but depending on the exact direction/project it might be more with Alex or more with Teun.
Theory of change: Soon, most important work will be done by AI. AI is going to increasingly advise people and help with important things, many of which are time-sensitive and path dependent, e.g., work on alignment/safety (including various things like how LLMs should behave given that they’re very persuasive); how to think about acausal trade; how to organize society. It seems good for AI to do well at those things.
Of course, a lot of the relevant skills for doing well at these tasks are the same skills that cause AI risk and that AI companies work on (and are incentivized to work on) by default; like coding, some kinds of forecasting, etc.
We want to make models better at things that are net positive for the future, but that likely won’t benefit much from said default training (or perhaps will even be made worse by such training – e.g., via sycophancy).
In practice, a lot of the tasks that we’re interested in from this perspective are what we call “conceptual”: tasks that are hard to verify and don't have clear ground truth but where we nonetheless feel like we can make progress through argument and reason.
You can visit conceptualreasoning.ai to get a sense of our work to date.
We also take a keen interest in projects directly aimed at making future acausal interactions go well.
None
I have two broad areas.
Security:
I am interested in building demonstrations for hacking real-world AI deployments to show that they are not secure. The goal is to force companies to invest in alignment techniques that can solve the underlying security issues.
Verification:
Verification via TEEs or ZKPs
I will meet 1-1 or as a group, depending on the interests as they relate to the projects. Slack communication outside of the 1-1.
I strongly prefer multiple short meetings over single long meetings, except at the start.
I'll help with research obstacles, including outside of meetings
For security:
You should have a strong security mindset, having demonstrated the willingness to be creative on this. I would like to see past demonstration of willingness to get your hands dirty and try many different systems.
For benchmarks:
As creative as possible, willingness to work on the nitty gritty, willingness to work really hard on problems other people find boring. Interests as far away from SF-related interests as possible.
Mentor(s) will talk through project ideas with scholar
This stream will focus on model motivations and character, open-ended environments, and new forms of misalignment.
This stream will focus on monitoring, stress-testing safety methods, and evals, with a focus on risks from scheming AIs. Examples include (black-box) AI control techniques, white-box monitors (probes etc.), chain-of-thought monitoring/faithfulness, building evaluation environments, and stress-testing mitigations.
For each project, we will have a weekly meeting to discuss the overall project direction and prioritize next steps for the upcoming week. On a day-to-day basis, you will discuss experiments and write code with other mentees on the project (though I'm available on Slack for quick feedback between meetings or to address things that are blocking you).
I structure the program around collaborative, team-based research projects. You will work in a small team, on a project from a predefined list. I organize the 12-week program into fast-paced research sprints designed to create and keep research velocity, so you should expect regular deadlines and milestones. I will provide a more detailed schedule and set of milestones at the beginning of the program.
I am looking for scholars with strong machine learning engineering skills, as well as a background in technical research. While I’ll provide weekly guidance on research, I expect scholars to be able to run experiments and decide on low-level details fairly independently most of the time. I’ll propose concrete projects to choose from, so you should not expect to work on your own research idea during MATS. I strongly encourage collaboration within the stream, so you should expect to work in teams of 2-3 scholars on a project, hence good communication and team skills are important.
We will most likely have a joint project selection phase, where we present a list of projects (with the option for scholars to iterate on them). Afterward, each project will have at least one main mentor, but we might also co-mentor some projects.
MATS 项目是一项为期 10 周的研究奖学金计划,旨在培养和支持从事人工智能对齐、透明度和安全领域工作的新兴研究人员。研究员将与世界一流的导师合作,获得专门的研究管理支持,并加入位于伯克利、致力于推动人工智能安全与可靠发展的活跃社区。该项目提供开展高影响力研究并开启人工智能安全领域长期职业生涯所需的架构、资源和指导。
MATS 导师均为来自人工智能安全、对齐、治理、领域建设及安全等广泛领域的顶尖研究人员。他们包括学术界人士、行业研究员以及独立专家,负责指导学者开展研究项目、提供反馈,并助力每位学者的研究成长。导师们的专业领域涵盖:
查看 往届及现任导师
关键日期
申请:
主项目将于 9 月 28 日至 12 月 4 日进行,获选研究员的延展阶段将于 12 月开始。
MATS 欢迎来自不同学术和专业背景的申请者——从机器学习、数学和计算机科学,到政策、经济学、物理学、认知科学、生物学和公共卫生,同时也欢迎没有传统研究背景的创业者、运营人员和领域建设者。主要要求是具备为人工智能安全做出贡献的强烈动机,并展现出技术能力、研究潜力或相关的运营经验。具备人工智能安全相关经验会有所帮助,但并非必要条件。