Keri Warr

Implementing SL4/5 and searching for differentially defense-favored security tools.

Stream overview

Developing techniques and technologies to enable SL4 and SL5 cybersecurity postures for LLMs, such as hardware and software supply chain management, confidential computing, weight exfiltration prevention, ML compute cluster security, and AI-powered insider threat detection.

Mentors

Keri Warr
Anthropic
,
Security Engineer
SF Bay Area
AI Systems Security

Keri is the technical lead of the Infrastructure Security Engineering Team at Anthropic, implementing SL4/5 and searching for differentially defense-favored security tools.

Read more

Mentorship style

I love asynchronous collaboration and I'm happy to provide frequent small directional feedback, or do thorough reviews of your work with a bit more lead time. A typical week should look like either trying out a new angle on a problem, or making meaningful progress towards productionizing an existing approach.

Fellows we are looking for

Essential:

  • Excited about doing Security Engineering work in particular
  • Non-zero experience in all three of: Software engineering, Cybersecurity, and Machine Learning
  • Ability to quickly unblock yourself
  • Willing to trade off against how shiny your project is in favor of doing work that will be impactful within the next 12 months

Preferred:

  • Significant depth in at least two of: Software engineering, Cybersecurity, or Machine Learning
  • Excited about delivering working open source code in addition to research papers

Can independently find collaboraters, but not required.

Project selection

Mentor(s) will talk through project ideas with scholar, or scholar will pick from a list of projects.

Streams

The Winter 2026 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

Oxford
Theory
AI Welfare
SF Bay Area
Control, Model Organisms, Scheming & Deception, Strategy & Forecasting
SF Bay Area
Theory
Interpretability
Tübingen
Dangerous Capability Evals, Agent Foundations, Adversarial Robustness, Monitoring, Scalable Oversight, Scheming & Deception
SF Bay Area
Policy and Governance
Dangerous Capability Evals, Policy & Governance
New York City
Monitoring, Dangerous Capability Evals, Scalable Oversight, Safeguards
SF Bay Area
Policy and Governance
Strategy & Forecasting, Policy & Governance
Montreal
Agent Foundations, Dangerous Capability Evals, Monitoring, Control, Red-Teaming, Scalable Oversight
SF Bay Area
Control, Model Organisms, Red-Teaming, Scheming & Deception