Nina Panickssery

Anthropic

Nina 参加了 2023 年夏季的 MATS 项目,并接受了 Evan Hubinger 的指导。在 MATS 项目期间,她发表了论文《Steering Llama 2 via Contrastive Activation Addition》,该论文荣获 ACL 2024 杰出论文奖。MATS 项目结束后,Nina 加入 Anthropic 担任研究科学家,并指导了多个致力于大语言模型对齐项目的 SPAR 和 MATS 小组。

The Summer 2023 cohort supported 60 fellows with 15 mentors, working across 12 different research areas. The program consisted of a remote 4-week training phase, an 8-week research phase in Berkeley, and a 4-month extension phase. MATS leadership co-founded the London Initiative for Safe AI (LISA) in September 2023 to provide a dedicated research space for AI safety researchers and organizations in London, and for MATS fellows to continue their research projects. Research projects were distributed across multiple areas, with approximately one-third focused on evaluations and capability demonstrations and one-fifth on mechanistic interpretability, alongside work on agent foundations, activation engineering, and cooperative AI.

参加 MATS 是快速提升人工智能安全研究技能、深入了解该领域并结识其他研究人员与合作伙伴的绝佳途径。此外,项目组精心设计的办公环境也极大地提高了工作效率。

Nina Panickssery