Victoria Krakovna

Conceptual research on deceptive alignment, designing realistic scheming propensity evaluations and honeypots. The stream will run in person in London, with scholars working together as a team.

Stream overview

Conceptual research on deceptive alignment, designing realistic scheming propensity evaluations and honeypots.

Mentors

Victoria Krakovna
Google DeepMind
,
Research scientist
London
Misalignment Science
AI Control and Monitoring
Capability and Propensity Evaluations

I am a research scientist on the AGI Safety & Alignment team at Google DeepMind. I focus on deceptive alignment and AI control, particularly [scheming propensity evaluations](https://arxiv.org/abs/2605.29729). My past research includes dangerous capability evals, power-seeking incentives, specification gaming, and avoiding side effects.

Read more

Mentorship style

During the program, we will meet once a week to go through any updates / results, and your plans for the next week. I'm also happy to comment on docs, respond on Slack, or have additional ad hoc meetings as needed.

Fellows we are looking for

- Collaborative software engineering: You are comfortable writing high-quality code independently and quickly. You are fluent in standard software engineering practices, e.g. version control. You can navigate a codebase written not entirely by you and can build on it productively. 

- Familiarity with deceptive alignment, safety cases, capability evaluations, AI control, and related topics. This makes it much easier for us to be on the same page about the goals of the project, and is important for making day-to-day project decisions in a conceptually sound way. 

- Conceptual research ability: You can come up with ideas for interesting experiments that align with the project plan. You prioritise well. 

- Good communication, team player: You can clearly communicate your research ideas, experimental methodology, results, etc, verbally or in writing (writing is more important). You can notice and express confusion, uncertainty, disagreement, or dissatisfaction and are willing to work through conflicts. You impartially consider other people’s ideas, disagree respectfully, can admit when you’re wrong, and are willing to commit to the team’s direction.

Project selection

I will talk through project ideas with scholars

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

London
Interpretability
London
Scheming & Deception, Dangerous Capability Evals, Control, Red-Teaming
SF Bay Area
Dangerous Capability Evals, Red-Teaming, Model Organisms, Control, Monitoring
Toronto
Interpretability
London
Control, Monitoring, Safeguards, Dangerous Capability Evals, Scheming & Deception
Chicago
Biorisk, Security, Safeguards
SF Bay Area
Interpretability, Agent Foundations
London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy & Governance, Strategy & Forecasting
Oxford
AI Welfare
SF Bay Area
Control, Model Organisms, Scheming & Deception, Strategy & Forecasting
SF Bay Area
Interpretability
Tübingen
Dangerous Capability Evals, Agent Foundations, Adversarial Robustness, Monitoring, Scalable Oversight, Scheming & Deception
SF Bay Area
Dangerous Capability Evals, Policy & Governance
New York City
Monitoring, Dangerous Capability Evals, Scalable Oversight, Safeguards