Stream overview

I'm interested in mentoring projects in several directions: 

  1. Safety or alignment pretraining: There is a large variety of topics in this direction, including ideas in generating synthetic pretraining data to help with improving robustness or alignment with human values, compliance with safety policies. 
  2. Data poisoning: As synthetic data becomes an increasingly large portion of pretraining corpora, new and more subtle forms of data poisoning become possible. I'm interested in projects that develop methods for detecting, characterizing, and defending novel forms of such attacks.
  3. The role of harmful data in building safer models: There is growing evidence that retaining some harmful content during pretraining, rather than filtering it all out, can improve a model's ability to reason about harms and ultimately become safer after post-training. I'd like to mentor work that deepens our understanding of this phenomenon: when does exposure to harmful data help vs. hurt, and how can we design pretraining pipelines that leverage this insight responsibly?

Mentors

Dylan Sam
OpenAI
,
Member of Technical Staff
SF Bay Area
AI Control and Monitoring
Alignment Training Methods

Dylan is a safety researcher at OpenAI, where he works on curating better/safer training data and monitoring models for harmful behavior.

Before that, he completed a PhD in the Machine Learning Department at CMU.

Read more

Mentorship style

Fellows we are looking for

Project selection

Streams

The Winter 2026 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

SF Bay Area
Dangerous Capability Evals
Boston
Policy and Governance
Adversarial Robustness, Policy & Governance, Red-Teaming, Safeguards
New York City
Control, Scalable Oversight, Red-Teaming, Model Organisms, Monitoring
SF Bay Area
Policy and Governance
Policy & Governance
SF Bay Area
Control, Monitoring, Dangerous Capability Evals
SF Bay Area
Security, Compute Infrastructure
London
Theory
Interpretability
London
Scheming & Deception, Dangerous Capability Evals, Control, Red-Teaming
SF Bay Area
Dangerous Capability Evals, Red-Teaming, Model Organisms, Control, Monitoring
Toronto
Interpretability
London
Control, Monitoring, Safeguards, Dangerous Capability Evals, Scheming & Deception
Chicago
Biorisk, Security, Safeguards
SF Bay Area
Interpretability, Agent Foundations
London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting