
参与 MATS 项目极大地提升了我的职业生涯、研究经验、人脉网络以及自信心。优秀的导师指导、高度的个人自主权、才华横溢且志向远大的社区,以及充足的资源支持,为“在实践中学习”并高效完成任务创造了绝佳条件。此外,我很难想象还有什么比 MATS 学者、导师以及整个 AI 安全社区正在解决的问题更重要的了。这不仅极具挑战性,而且意义深远。在这个领域,大家怀揣着共同的信念,结下了深厚的友谊。加入我们吧!
Naci 的工作专注于人工智能开发与应用中的透明度及验证机制。这些机制旨在推动国际社会就人工智能的克制与审慎达成共识,并实现对人工智能技术及相关利益方的民主监督。Naci 拥有亚琛工业大学物理学硕士学位,曾在 SPAR 的 Aaron Scher 和 MATS 的 Mauricio Baker 指导下,从事人工智能硬件技术及供应链方面的研究。

There's life pre-MATS and life post-MATS. It was the inflection point that set me up to become a technical AI safety researcher. I don't think there are other opportunities as good at getting early-career people integrated into AI safety. The in-person program was the most impactful and high-energy two months I've ever been a part of, and it's my number one recommendation to people considering work on AI safety.
Jesse Hoogland is the executive director of Timaeus, an AI safety research organization studying developmental interpretability and singular learning theory. He was a MATS scholar during MATS 3.0 and 3.1 in Evan Hubinger's Deceptive AI stream. During this period, he became interested in understanding how AI systems develop during training. This led to him helping to organize the SLT and Alignment conference and the DevInterp conference, which resulted in the developmental interpretability research agenda.

Apollo almost certainly would not have happened without MATS. One of the core reasons why starting an organization is hard is because the founding members need to know and trust each other. It is often hard to find people with similar agendas that you also personally enjoy working with in a systematic manner. MATS implicitly created such an environment because it enabled many of us to understand what everyone else is working on, get to know them personally and see their research progress without having to commit to anything in particular.
Marius took part in MATS Winter 2022/23 Cohort under the mentorship of Evan Hubinger (Anthropic). He published multiple pieces on mechanistic interpretability on LessWrong including work on maximum data dimension and double descent. He is currently the CEO and Director of Apollo Research, a new London-based technical alignment organization. Previously, he did a Ph.D. in Machine Learning and conducted independent alignment research. Read more on his website.

在参加 MATS 之前,我对人工智能对齐领域有着浓厚的兴趣,但缺乏前沿研究的相关技能,也不知从何入手。直接得益于 MATS,我实现了以下目标:(1) 对人工智能安全领域最核心的问题及其相关社区的结构有了相对完整的理解;(2) 产出了清晰且具有重要意义的研究成果,这让我有信心全职投身于该领域;(3) 结识了广泛的现任及未来合作者,他们带来了极其多元的视角。关于第三点,MATS 汇聚的人才令人惊叹,他们解决问题的动力极其强烈。如果未来人工智能对齐这一宏大工程最终取得成功,届时人们会发现,关键问题与解决方案中超过两位数百分比的贡献都归功于 MATS 的校友,对此我一点也不会感到惊讶。
我是一名独立人工智能安全研究员,目前专注于机械可解释性和训练过程透明度。

MATS 是快速提升技能并建立人工智能安全领域人脉的最佳途径,我强烈推荐。
Joseph Bloom 是 英国人工智能安全研究所(UK AI Security Institute)的模型透明度负责人,致力于研究失控风险、可监控性和可解释性之间的交叉领域。他的团队近期发表了关于 针对“藏拙”行为(Sandbagging)的游戏审计的研究。Joseph 曾是 MATS 5.0 计划中 Neel Nanda 的学员。他此前曾担任 TransformerLens 软件包的维护者,开发了 SAE Lens 软件包,并以 LASR 导师身份发表了 A is for Absorption 。Joseph 拥有墨尔本大学计算生物学与统计学双学位。

MATS 帮助我提升对齐领域技能的速度,比我原本通过自学基础贝叶斯主义(infra-bayesianism)快了三倍多。当时我只是因为喜欢数学而自学,对对齐领域中哪些部分至关重要并没有深刻的见解。MATS 让我对对齐问题有了更深层次的认识,此后我能够专注于解决问题的核心,并理清了自己认知中最主要的困惑。
Thomas 参加了 John Wentworth 指导的 2022 年夏季班和 Nate Soares 指导的 2023 年冬季班。在此期间,他撰写了一份关于 AI 安全研究方法的详细综述。随后,他在 MIRI 继续开展 SERI MATS 的研究工作,之后离职创办了 AI 安全倡导组织——人工智能政策中心(Center for AI Policy)。目前,他是 AI 未来项目(AI Futures Project)的研究员,同时担任 LTFF 的客座基金经理。


参加 MATS 是快速提升人工智能安全研究技能、深入了解该领域并结识其他研究人员与合作伙伴的绝佳途径。此外,项目组精心设计的办公环境也极大地提高了工作效率。
Nina 参加了 2023 年夏季的 MATS 项目,并接受了 Evan Hubinger 的指导。在 MATS 项目期间,她发表了论文《Steering Llama 2 via Contrastive Activation Addition》,该论文荣获 ACL 2024 杰出论文奖。MATS 项目结束后,Nina 加入 Anthropic 担任研究科学家,并指导了多个致力于大语言模型对齐项目的 SPAR 和 MATS 小组。

我强烈推荐 MATS!对于那些希望投身人工智能安全技术研究的人来说,MATS 是我的首选。我在 MATS 获得的指导和融入的社区氛围,不仅让我作为研究人员迅速成长,也为我探索有价值的研究方向提供了广阔空间。
Cody Rushing 是德克萨斯大学奥斯汀分校计算机科学专业的本科生。他目前正与 Buck Shlegeris 及 Redwood Research 合作开展人工智能控制方面的研究,并将于秋季继续这项工作。
https://starship006.github.io/

MATS was a life changing experience. I met and got mentored by amazing people, and I learned so much in such a small amount of time. Looking back at me before this program, I don't think I could even recognize myself 8 month ago. Even though I have no academic background, I felt listened, empowered and supported in order to tackle the biggest challenges that I (and possibly we) have ever faced.
After MATS, I worked as a contractor for METR evaluating GPT-4 pre-release. I then co-founded PRISM Eval and created an automated red-teaming system (BehaviorElicitiationTool: https://github.com/qfeuilla/BehaviorEliciationTool) that I presented at the Paris AI Summit. I am now founding WeaveMind (https://weavemind.ai/) at Seldon Lab Batch 2.


MATS helped me get deeper into AI safety research by motivating me to get up to speed with current research and giving me access to mentorship from an expert in AI safety, as well as a smart and talented cohort and a large network of researchers. It also provided infrastructure such as office space in Berkeley and a generous stipend. SERI MATS worked as a matchmaker between Evan Hubinger and me and thus helped me get involved in his projects, which would have been harder to do otherwise. I feel like I have developed faster as a researcher since doing MATS.
Johannes completed the MATS Summer 2022 Cohort under the mentorship of Evan Hubinger (then a Research Fellow at MIRI). As a result of MATS, Johannes co-authored the paper Conditioning Predictive Models: Risks and Strategies with Evan as a lead author. He also published a follow-up paper on Incentivizing honest performative predictions with proper scoring rules at the UAI 2023 conference. After MATS, Johannes started a PhD in Computer Science at CHAI. Since 2024, he Johannes has been working at Anthropic on alignment stress-testing.

Working in a team environment, particularly one as stimulating as MATS, was a transformative experience. It not only refined my research skills but also instilled a newfound entrepreneurial spirit in me. The program encouraged me to think beyond the conventional, to innovate, and to take risks. Additionally, the array of skills I acquired during my time at MATS was vast. I delved deep into research engineering, honed my science communication abilities, and even tapped into the art of fundraising. These skills, I believe, are indispensable and have equipped me to navigate the ever-evolving world of research with confidence. In conclusion, I wholeheartedly endorse the MATS program. To anyone considering embarking on this journey, you are not only signing up for an unparalleled research experience but also a lifetime of growth, learning, and camaraderie.
I'm working on AI Safety Connect, a new organization convening diplomatic and AI Safety stakeholders at the highest level - think UN, India Impact Summit etc. We are also seeding a few other projects, like engaging the UAE in AI Safety and helping prevent critical coordination failures among frontier labs.


MATS was an excellent environment to get productive work done and a fantastic resource to improve my future impact in AI alignment. I made connections, learned a great deal about my mentor's subfield and alignment in general, and was fired up to keep working when I got back to Australia. Since MATS I've been funded for a project with a collaborator I met at MATS, and gotten significantly further in the hiring process for orgs than before.
Previous UK AISI employee experienced in frontier LLM evaluation, now looking to contribute to technical AI safety and reducing extinction risks from misaligned AGI systems.


Ethan spent a lot of time discussing our research with us and gave great advice on direction. He unblocked us in various ways, such as getting access to more models or to lots of compute budget. He connected us with lots of great people, some of whom became collaborators. And he was a very inspiring mentor to work with.
Dan Valentine is a Member of Technical Staff at Anthropic, an AI safety and research company. His work is primarily focused on AI safety and alignment research, including scalable oversight methods and understanding how AI models interact with data and prompts.
自 2021 年底以来,已有 631 名研究人员参与了 MATS 项目,产出了 220 多篇研究论文,并加入顶尖 AI 实验室,或创立了推动 AI 对齐、透明度和安全性进步的新组织。
10% 的 校友 在 MATS 期间或之后共同创立了人工智能安全组织和研究团队。
MATS 校友创立的组织包括 Aether、AI Safety Argentina、Algoverse AI Safety Fellowship、Apollo Research、ARENA、Athena、Cadenza Labs、Catalyze Impact、Center for AI Policy、Contramont Research、Coordinal Research、Decode Research、Dovetail Research、Freestyle Research、Fulcrum、General Analysis、GPAI Policy Lab、Groundless、Leap Labs、LISA、Luthien Research、MIRI Technical Governance Team、Poseidon Research、Principled Agents、PRISM Eval、Queensland AI Safety Initiative、Reciprocal Research、Resolution、Robocurve、Simplex、Security Level 5、StakeOut AI、Timaeus、Theorem Labs、WeaveMind 和 Workshop Labs。
75% 的 校友 目前正从事人工智能对齐、透明度和安全领域的工作。
MATS 校友已被 Anthropic、Google DeepMind、OpenAI、Meta AI、UK AISI、Redwood Research、METR、RAND CAST、Coefficient Giving、ARC、FAR.AI、Apollo Research、Truthful AI、Goodfire、LawZero、MIRI、CAIF、Center on Long-Term Risk、Beneficial AI Foundation、SaferAI、Haize Labs、EleutherAI、Harmony Intelligence、Conjecture 等领先组织聘用,并加入了 UC Berkeley CHAI、NYU ARG、NU Bau Lab 和 Mila 等学术研究小组。
MATS 提供导师指导、研究资金、住宿和社区支持,让研究人员能够全身心投入到解决全球最重要的问题中。
研究员每周可获得 1250 美元的生活津贴,用于支付相关费用。
研究员在项目全程可获得餐饮供应,并享有伯克利的住宿安排,以及伦敦的住房补贴。
研究员每周可获得 2000 美元的计算资源额度,用于支持实验和评估工作。
研究员可以使用位于伯克利和伦敦的办公空间,并能与同行研究人员进行日常协作。
研究员将参加由人工智能对齐领域专家主持的研讨会、工作坊和客座讲座。
研究员将与专属研究经理合作,由其协助规划项目、把控进度并解决障碍。
研究员将获得来自 AI 对齐、治理和安全领域顶尖研究人员的指导。
研究员可申请加入 MATS 延期项目,进行为期 6 至 12 个月的额外资助研究;超过 80% 的研究员会获得录取。
研究员可以建立联系,并获得在更广泛的 AI 对齐生态系统中进行交流的机会。
MATS 研究员产出的研究成果涵盖了推进 AI 安全、韧性和理解力的各个方面。研究员们通过机械可解释性、稀疏特征分析、潜在表征研究及其他技术,深入探究现代 AI 系统的内部运作机制。
Towards Understanding Sycophancy in Language Models
Human feedback is commonly utilized to finetune AI assistants. But human feedback may also encourage model responses that match user beliefs over truthful ones, a behaviour known as sycophancy. We investigate the prevalence of sycophancy in models whose finetuning procedure made use of human feedback, and the potential role of human preference judgments in such behavior. We first demonstrate that five state-of-the-art AI assistants consistently exhibit sycophancy across four varied free-form text-generation tasks. To understand if human preferences drive this broadly observed behavior, we analyze existing human preference data. We find that when a response matches a user's views, it is more likely to be preferred. Moreover, both humans and preference models (PMs) prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time. Optimizing model outputs against PMs also sometimes sacrifices truthfulness in favor of sycophancy. Overall, our results indicate that sycophancy is a general behavior of state-of-the-art AI assistants, likely driven in part by human preference judgments favoring sycophantic responses.
作者:
Meg Tong
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, Ethan Perez
日期:
Oct 20, 2023
Stealing Reasoning Traces from Proprietary LLM APIs
Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem. We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly. This vulnerability enables four distinct attack vectors. First, it circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model's reasoning, as we demonstrate across Anthropic, OpenAI, and Google. Second, it allows for large-scale private data extraction. Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials. Third, it inadvertently reveals hazardous information hidden within the reasoning process, even in cases where the model's final, visible output safely rejects a malicious request. Fourth, attackers can leverage this flaw to execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Following responsible disclosure, we propose concrete cryptographic and system-level mitigations to secure client-side reasoning.
作者:
Alexander Panfilov, Joachim Schaeffer
Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, Maksym Andriushchenko
日期:
Aug 10, 2026
点击相应方向以了解更多关于申请流程、申请人画像及重点领域的信息。
.webp)
MATS 旨在解决人工智能安全领域的人才瓶颈,并为我们认为当今世界最紧迫且最缺乏人才的问题培养专业人员:降低未对齐人工智能带来的风险。我们坚信,来自不同背景的有志研究人员都有潜力为对齐研究领域做出实质性贡献。我们致力于提供必要的培训、后勤支持和社区环境,以助力这一转型。我们还为研究员提供经济支持,确保他们的生活稳定与安全。
MATS Research 是一家独立的 501(c)(3) 公益慈善机构(EIN: 99-0648563)。