MATS Fellow:
Axel Højmark, Govind Pimpale, Arjun Panickssery
Authors:
Axel Højmark, Govind Pimpale, Arjun Panickssery, Marius Hobbhahn, Jérémy Scheurer
Citations
Abstract:
要降低 AI 系统带来的风险,我们需要准确评估其能力,尤其是在系统很少展示某项能力时,这项工作会格外困难。Phuong 等人提出了两种方法,旨在更准确地估计 AI 智能体成功完成某项任务的概率。里程碑方法将任务拆分为子任务,以改进整体成功率估计;专家 best-of-N 方法则以人工指导作为模型独立表现的近似指标。我们将这两种方法作为蒙特卡洛估计量进行分析,发现它们虽然都比朴素蒙特卡洛抽样更能降低方差,却也会引入偏差。实验结果表明,由于假设过于严格,里程碑方法会低估许多真实任务的实际解题率。专家 best-of-N 方法在所有任务上都低估得更严重,原因是其重加权因子存在根本缺陷。为提高对智能体困难任务能力估计的准确性,我们建议后续研究借鉴蒙特卡洛估计量领域的丰富文献。
Synthetic Persona Pretraining: Alignment from Token Zero
Authors:
Julian Minder
Date:
August 13, 2026
Citations:
The MATS Program is an independent research and educational initiative connecting emerging researchers with mentors in AI alignment, governance, and security.
Each MATS cohort runs for 12 weeks in Berkeley, California, followed by an optional 6–12 month extension in London for selected scholars.