Joseph Bloom

英国人工智能安全研究所

Joseph Bloom 是 英国人工智能安全研究所(UK AI Security Institute)的模型透明度负责人,致力于研究失控风险、可监控性和可解释性之间的交叉领域。他的团队近期发表了关于 针对“藏拙”行为(Sandbagging)的游戏审计的研究。Joseph 曾是 MATS 5.0 计划中 Neel Nanda 的学员。他此前曾担任 TransformerLens 软件包的维护者,开发了 SAE Lens 软件包,并以 LASR 导师身份发表了 A is for Absorption 。Joseph 拥有墨尔本大学计算生物学与统计学双学位。

The Summer 2024 cohort marked a significant expansion, supporting approximately 90 fellows with 40 mentors—the broadest mentor selection in MATS history. This cohort incorporated MATS as a 501(c)(3) nonprofit organization, formalizing its institutional structure. The program expanded its research portfolio to include at least four governance mentors alongside technical research streams, reflecting growing interest in AI policy and technical governance work. The 10-week research phase continued in Berkeley, with fellows conducting work across mechanistic interpretability, evaluations, scalable oversight, and governance research. Notable outputs from this cohort include research on targeted manipulation and deception in LLMs trained on user feedback, which was accepted to NeurIPS workshops, and contributions to an AI safety via debate paper that won best paper at ICML 2024. One fellow co-founded Decode Research, a new AI safety organization focused on building interpretability tools.

MATS 是快速提升技能并建立人工智能安全领域人脉的最佳途径,我强烈推荐。

Joseph Bloom