Arthur Conmy

Anthropic

Research Engineer

Links

Focus

AI Control and Monitoring, Alignment Training Methods, Interpretability

Arthur Conmy is a Member of Technical Staff at Anthropic. His interests are in automating interpretabilityfinding circuits and making model internals techniques useful for AI Safetyparticularly with Sparse Autoencoders. Previously, he worked at Google DeepMind and Redwood Research (and did the MATS Program!).