This stream focuses on building a Science of Scheming, i.e. what are the mechanisms by which future models might become schemers, even though current models are not. We want to discover empirical Scaling Trends for Scheming. For example, does deceptive alignment become easier to discover with improved model capabilities?
This stream focuses on building a Science of Scheming, i.e. what are the mechanisms by which future models might become schemers, even though current models are not. We want to discover empirical Scaling Trends for Scheming.
Some example questions:
Does deceptive alignment become easier to discover with improved model capabilities?
When doing alignment training, a model can achieve low loss either via internalizing the rewarded values (inner alignment) or via pretending to internalize the rewarded values in order to defect when oversight lapses (deceptive alignment). Which one of these happens is a complicated question of inductive bias. As far as anyone knows, deceptive alignment has only arisen under strongly nudged settings (e.g. sleeper agents paper or opus-3 alignment faking).
Intuitively, we expect that dumb models are extremely unlikely to naturally converge to deceptive alignment without extreme handholding, but the more capable a model gets of situational awareness and of long-term goal-pursuit, the easier it might be for the model to discover deceptive alignment. Can we quantify the degree of "handholding" that is needed before a model becomes deceptively aligned, and then measure whether the amount of handholding decreases with additional model capabilities? If so, it could serve to forecast if and when we should expect deceptive alignment to emerge in the wild.
This project would involve finetuning and RL'ing a bunch of model organisms, as well as finding creative ways to quantify "degree of handholding".
What drives the emergence of reward-seeking?
One of the precursors for deceptive alignment is "reward-seeking", i.e. models actively thinking about what would be rewarded and then optimizing for this. It appears that current frontier models are fairly well-described as reward-seekers, but it is far from obvious why and how this tendency has emerged. In particular, there is a distinction between "seeking reward" and "taking actions that lead to high reward". Determining whether a given training pipeline should be expected to create a reward-seeker is again a complex question of inductive bias and training dynamics.
We hypothesize that reward-seeking emerges because, on average, it is useful: rollouts where the model actively thinks about and pursues reward tend to perform slightly better than rollouts where it doesn't. Over sufficiently high-compute RL, this slowly drives up reward-seeking cognition. We want to test this hypothesis by creating toy settings where we can control how useful reward-seeking reasoning is for reward. The hope is that understanding the emergence of reward-seeking will also teach us about the emergence of deceptive alignment.
This project involves running a large number of RL experiments and deeply analyzing the RL dynamics. Practical experience with running RL on LLMs is a definite plus.
Standard (1-2 hours of weekly 1:1s)
London
Indifferent
Indifferent
Teun leads the RL dynamics project at Apollo Research. The RL dynamics project fits under the umbrella of our science of scheming approach. The current main focus is to understand reward-seeking dynamics during reinforcement learning, for which we combine theory and empirics.
Before that, he worked on control, sandbagging, building an AI superforecaster, and more. He took part in MATS 5.0!
Teun is also a board member for ENAIS and SAIN.
Alexander Meinke is Head of Research at Apollo Research. His team empirically studies how "scheming" can emerge in future AI systems.
He started working on AI safety research in 2023 with Owain Evans in the MATS 4.0 cohort.
Before that, he completed his PhD on adversarial robustness at the University of Tübingen, Germany. He holds a B.Sc. and M.Sc. in Physics.
Ideal candidates would have (some of):
We will set the high-level project direction, as described above. It's not fully clear what exactly the project will look like by the time you start in September. All projects will be in the direction of the Science of Scheming post.
You’d work with the two of us, but depending on the exact direction/project it might be more with Alex or more with Teun.