MATS mentors are advancing the frontiers of AI alignment, transparency, and security

Stephen "Cas" Casper is a computer scientist and an Assistant Professor of Public Policy at the Harvard Kennedy School and a Faculty Affiliate of the Harvard School of Engineering and Applied Sciences. Prior to joining Harvard, he completed his PhD at MIT and did a research residency with the UK AI Security Institute. He is a writer for the International AI Safety Report and a lead writer for the Singapore Consensus. His research has been recognized with a Hoopes Prize, an ML Safety Workshop best paper award, a BioSafeGenAI best paper runner-up, a GenLaw spotlight paper award, a TMLR outstanding paper finalist distinction, and a handful of mentions in news articles and newsletters. Find him on Google ScholarTwitter (sorry), BlueSky, and LinkedIn.

Focus:
Policy and Governance
Adversarial Robustness and Safeguards, Alignment Training Methods, Policy and Governance, Technical AI Governance

I am a research scientist on the AGI Safety & Alignment team at Google DeepMind. I focus on deceptive alignment and AI control, particularly scheming propensity evaluations. My past research includes dangerous capability evals, power-seeking incentives, specification gaming, and avoiding side effects.

Focus:
Empirical
AI Control and Monitoring, Capability and Propensity Evaluations, Misalignment Science
Mauricio Baker
RAND
,
Technical AI Policy Research Scientist; DPhil (PhD) student

Mauricio researches AI policy at RAND and Oxford. His work has focused on verification of international agreements on AI. He’s more broadly interested in technical AI governance. Previously, Mauricio contracted with OpenAI and did a master's in Computer Science at Stanford University.

Focus:
Systems Security
Policy and Governance, Technical AI Governance, AI Systems Security
Alex Mallen
Redwood Research
,
Member of Technical Staff

Alex is a member of technical staff at Redwood Research.

Focus:
Empirical
AI Control and Monitoring, Misalignment Science, Forecasting and Strategy, Interpretability, Theoretical Alignment and Formal Methods
Abram Demski
AFFINE
,
Research Scientist

Abram Demski is an AI Safety researcher specializing in Agent Foundations, best known for Embedded Agency (co-written with Scott Garrabrant). His overall approach primarily involves deconfusion research in relation to various concepts related to AI risks, including agency, optimization, trust, meaning, understanding, interpretability, and computational uncertainty (more commonly but less precisely known as bounded rationality). More specifically, his recent work focuses on modeling trust, with the objective of clarifying conditions under which humans can justifiably trust AI.

Focus:
Theory
Agent Foundations, Theoretical Alignment and Formal Methods

Paul Riechers is a researcher and scientific leader with deep expertise in the physics of information and the fundamental limits of learning and prediction.  He co-founded the Simplex AI safety research organization with Dr. Adam Shai, applying insights from theoretical physics and neuroscience to build foundational understanding of internal representations and emergent behavior in neural networks.  Paul earned a PhD in theoretical physics and an MS in electrical and computer engineering from UC Davis. Prior to founding Simplex, he spent five years as a Research Fellow at Nanyang Technological University in Singapore. He is also co-founder of the Beyond Institute for Theoretical Science (BITS), a former Senior Fellow at UCLA’s Mathematics of Intelligences program at IPAM, and has served as both a MATS scholar and mentor. Paul has co-organized multiple workshops on AI interpretability and alignment, and now co-leads the growing Simplex team with support from the Astera Institute.

Focus:
Empirical
Interpretability
James Lucassen
Redwood Research
,
Member of Technical Staff

James is a member of technical staff at Redwood Research.

Focus:
Empirical
AI Control and Monitoring, Misalignment Science
Aryan Bhatt
Redwood Research
,
Member of Technical Staff

Aryan is a senior member of technical staff at Redwood Research.

Focus:
Empirical
AI Control and Monitoring, Misalignment Science
Jack Lindsey
Anthropic
,
Member of Technical Staff

Hi, I'm Jack! I'm interested in understanding the cognition of modern language models, so that we can make them more reliable and aligned with human values. Currently, I lead the "Model Psych" team at Anthropic. We study the internal basis of higher-level cognitive phenomena in LLMs, like introspection, situational awareness, personas, and representations of emotion. We apply these techniques to audit Anthropic’s production models, for instance by monitoring their neural activity for signatures of deception, manipulation, or awareness of being evaluated. Previously, I did my PhD in the Center for Theoretical Neuroscience at Columbia University. For a list of my publications, see my Google Scholar profile.

Focus:
Empirical
AI Control and Monitoring, Misalignment Science, Interpretability
Vivek Hebbar
Redwood Research
,
Member if Technical Staff

Vivek is a member of technical staff at Redwood Research.

Focus:
Empirical
AI Control and Monitoring, Misalignment Science

Frequently asked questions

What is the MATS Program?
Who are the MATS Mentors?
What are the key dates of the MATS Program?
Who is eligible to apply?
How does the application and mentor selection process work?