Andrew Saxe & Nischal Mainali

This stream builds theories of deep learning phenomena from simplified models, covering emergent misalignment, unlearning under fine-tuning, and reward hacking and other RLVR pathologies through an extension of the RL perceptron.

Stream overview

I'm open to fellows picking their own projects focused on learning dynamics and AI safety. There's no shortage of interesting problems in this area at present.

​

Below are some concrete ideas which I have a rough sketch of how to tackle:

​

-Theories of emergent misalignment based on simplified models of deep learning  

​

-Theories of unlearning based on simplified models of fine tuning 

​

-Theories of reward hacking and other RLVR pathologies by extending the RL perceptron 

​

These projects would all require building on existing theoretical models in the statistical physics/nonlinear dynamics tradition to obtain insight into these phenomena. This requires a mix of mathematical derivations, simulations to verify insights in simplified models hold in more complex models, and conceptual work to ensure the idealized system is an important model to understand.

Mentorship style:

Standard (1-2 hours of weekly 1:1s)

Location during program:

London

London location preference:

Weak preference

Berkeley location preference:

No preference for this location

Mentors

Andrew Saxe
Principa
,
Professor/Scientific Director
No items found.

Andrew Saxe is Professor of Theoretical Neuroscience and Machine Learning at University College London, and a joint group leader at the Gatsby Computational Neuroscience Unit and the Sainsbury Wellcome Centre. Saxe runs the Theory of Learning Lab, which works on the theory of learning in brains and deep networks.

Read more
Nischal Mainali (Nisch)
Principa
,
Research scientist
No items found.

Nischal is a research scientist at Principia interested in theoretical understanding of AI systems and using that understanding for interpretability, i.e., theory-first interpretability. He is currently interested in solvable models of learning dynamics and their applications, and mean-field theories of learning and representation formation, and would be interested in either pushing the theory frontier and developing basic understanding of safety-relevant phenomena (e.g., silent alignment, representation structure in mean-field networks, solvable models of superposition, etc.) or applying various ideas coming from theory for practical purposes (early detection of sudden changes in NNs during learning, new weight space/representation space interpretability tools, etc.).

Read more

Fellows we are looking for

Math background covering statistical physics tools or nonlinear dynamical systems

Project selection