Arthur Conmy

Arthur Conmy's MATS Stream focuses on evaluating interpretability techniques on current and future AI Safety problems.

This can involve creating new safety techniques, as well as creating benchmarks and measuring performance against baseline techniques.

Stream overview

I am broadly interested in research directions scholars are excited about, that can advance the quality of our AI Safety tools, and our confidence in them. Three particular areas of research that seem promising to me are:

  1. Reasoning Model Interpretability
  2. Eliciting Strange Behaviors from Models 'in the wild' (i.e. not by training model organisms)
  3. Model Organisms too :)

Mentors

Arthur Conmy
Anthropic
,
Research Engineer
London
Interpretability
AI Control and Monitoring
Alignment Training Methods

Arthur Conmy is a Member of Technical Staff at Anthropic. His interests are in automating interpretabilityfinding circuits and making model internals techniques useful for AI Safetyparticularly with Sparse Autoencoders. Previously, he worked at Google DeepMind and Redwood Research (and did the MATS Program!).

Read more

Mentorship style

I meet 1h/week, in group meetings (scheduled).

I also fairly frequently schedule ad hoc meetings with scholars to check on how they're doing and to address issues or opportunities that aren't directly related to the project.

I'll help with research obstacles, including outside of meetings.

Fellows we are looking for

Executing fast on projects is highly important. But also having a good sense of which next steps are correct is also valuable, though I enjoy being pretty involved in projects, so it's somewhat easier for me to steer projects than it is for me to teach you how to execute fast from scratch. It helps to be motivated to make interpretability useful, and use it for AI Safety, too.

I will also be interviewing folks doing Neel Nanda's MATS research sprint who Neel doesn't get to work with.

I think collaborations are strong, so I would try to pair you with, for example, some of my MATS extension scholars, or most likely other scholars I take this round.

Project selection

Mentor(s) will talk through project ideas with scholar.

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting
Oxford
Theory
AI Welfare
SF Bay Area
Control, Model Organisms, Scheming & Deception, Strategy & Forecasting
SF Bay Area
Theory
Interpretability
Tübingen
Dangerous Capability Evals, Agent Foundations, Adversarial Robustness, Monitoring, Scalable Oversight, Scheming & Deception
SF Bay Area
Policy and Governance
Dangerous Capability Evals, Policy & Governance
New York City
Monitoring, Dangerous Capability Evals, Scalable Oversight, Safeguards
SF Bay Area
Policy and Governance
Strategy & Forecasting, Policy & Governance
Montreal
Agent Foundations, Dangerous Capability Evals, Monitoring, Control, Red-Teaming, Scalable Oversight
SF Bay Area
Control, Model Organisms, Red-Teaming, Scheming & Deception