He He

I'm interested in mentoring projects related to reward hacking and monitoring (agentic) models that produce long and complex trajectories. Scholar will have freedom to propose projects within this scope. Expect 30-60min 1-1 time on zoom.

Stream overview

I'm interested in mentoring projects related to reward hacking and monitoring (agentic) models that produce long and complex trajectories.

Location during program:

New York City

Mentors

He He
New York University
,
Associate Professor
Misalignment Science
AI Control and Monitoring
Capability and Propensity Evaluations
对抗鲁棒性与安全防护

He He is an associate professor at New York University. She is interested in how large language models work and potential risks of this technology.

Read more

Fellows we are looking for

  • Strong engineering skills and experience in training deep learning models
  • Familiarity with modern large scale RL pipelines (e.g., using frameworks such as verl)

Project selection

Week 1-2: Mentor will provide high level directions or problems to work on, and scholar will have the freedom to propose specific projects and discuss with mentor.

Week 3: Figure out detailed plan of the project.