I work on adversarial stress-testing at METR. Fellows in this stream will build model organisms of misalignment that require real cognition to decide whether to misbehave, then test which detection methods actually catch them, with the goal of turning whatever works into an alignment evaluation we can run on real models.
I work on adversarial stress-testing at METR. The projects I am most excited about right now:
Sydney works on adversarial stress-testing at METR. She studied computational biology at Stanford and co-founded the Atlas Fellowship.
The Winter 2026 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.