MATS Fellow:
Wilson Wu, Louis Jaburi, Jacob Drori
Authors:
Wilson Wu, Louis Jaburi, Jacob Drori, Jason Gross
Citations
Abstract:
近期一系列机械可解释性研究,着重反向工程在有限群二元运算任务上训练的神经网络所执行的计算。我们研究了在该任务上训练的单隐藏层神经网络内部机制,揭示此前未识别的结构,并更完整地描述这类模型,为统一以往研究的解释迈进一步(Chughtai 等,2023;Stander 等,2024)。值得注意的是,这些模型会分别对每个输入参数近似满足等变性。我们将解释转化为模型性能的紧凑证明,以验证它是否适用于大量此类网络;这一方法可以定量评估我们对模型内部机制的解释是否忠实且简洁。正文主要讨论对称群 S5。对于在该群上训练的模型,我们的解释可给出模型准确率保证,运行速度比穷举快 3 倍,并为我们训练的 45% 模型给出至少 95% 的准确率下界。仅使用以往研究的解释,我们无法得到非平凡且非空泛的准确率界。
Synthetic Persona Pretraining: Alignment from Token Zero
Authors:
Julian Minder
Date:
August 13, 2026
Citations:
The MATS Program is an independent research and educational initiative connecting emerging researchers with mentors in AI alignment, governance, and security.
Each MATS cohort runs for 12 weeks in Berkeley, California, followed by an optional 6–12 month extension in London for selected scholars.