MATS Fellow:
Joachim Schaeffer, Alexander Panfilov
Authors:
Joachim Schaeffer, Alexander Panfilov, Thomas Jiralerspong, Guillaume Lajoie, Jonas Geiping, Yoshua Bengio, Roland S. Zimmermann
Citations
Abstract:
AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untrusted model's trajectory. If the acting model detects such an intervention, it may infer properties of the monitor and adapt to evade the control protocol. We introduce CIAware-Bench, a benchmark for measuring control intervention (CI) awareness across frontier models. CIAware-Bench tests whether models can distinguish their own trajectories from those modified by a control intervention. The benchmark comprises four task domains (essay writing, BigCodeBench, Bash Arena, and SHADE-Arena), with options to vary trajectory watermarking, side-task presence, and the control protocol. Evaluating thirteen acting models with eight intervening models shows substantial variation between settings and model pairs. CI awareness rises sharply for GPT-6 Astra and the Claude 5 models (Fable 5 and Opus 5). When explicitly probed, Astra reaches mean AUROC of 0.90 on Essay, 0.91 on BigCodeBench, 0.86 on Bash Arena and 0.77 on SHADE-Arena. Fable 5 and Opus 5 both reach 0.77 on Essay, with less consistent gains in the other environments. On SHADE-Arena, we demonstrate that telling an acting model that an action was replaced and asking it to adapt leads to moderate improvements in monitor evasion rates. In summary, our results highlight that control evaluations should already assume perfect CI awareness for conservative safety estimates, and that protocol design should explore countermeasures that make interventions harder to detect.
Counterfactual Debugging the World Model Transfer Gap
Authors:
Mingxuan Li
Date:
December 8, 2026
Citations:
CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs
Authors:
Joachim Schaeffer, Alexander Panfilov
Date:
October 5, 2026
Citations:
The MATS Program is an independent research and educational initiative connecting emerging researchers with mentors in AI alignment, governance, and security.
Each MATS cohort runs for 12 weeks in Berkeley, California, followed by an optional 6–12 month extension in London for selected scholars.