Peter Henderson’s stream focuses on developing safe, aligned AI agents, with projects on scalable oversight rules informed by law and game theory, safe long-horizon exploration, and measuring “jagged” capability/safety frontiers. Scholars will join an independently driven, engineering-heavy research environment, collaborating with other MATS scholars and PhD students, with weekly 1:1s and active async mentorship.
I'd be interested in a variety of potential projects, but three directions that I'm excited by are:
Peter is an assistant professor at Princeton University, where he works on reinforcement learning, alignment, and law. He received a J.D. and Ph.D. in computer science from Stanford University.
45 min weekly meetings by default for high-level guidance. I'm active on Slack for quick questions or conceptual (not code) debugging. Expect async back-and-forth on experiment design and results between meetings. Scholars can also schedule ad-hoc calls if they're stuck or want to brainstorm—just ping me on Slack. Other team members (PhD students) will also be around to help brainstorm, getting unstuck.
Essential:
Nice to have, but not necessary:
Not a good fit:
Collaborators will be other MATS scholars, as well as PhD students/post-docs at multiple institutions (including Princeton).
Mentors in the group will pitch projects, and scholars will try ones they find interesting for a week. We'll iterate together at the end of week 1 and pick final assignments in week 2.
The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.