Understanding what AI systems are thinking seems important for ensuring their safety, especially as they become more capable. For some dangerous capabilities like deception, it’s likely one of the few ways we can get safety assurances in high stakes settings.
The ability to reverse engineer what neural networks have learned promises one of the few ways that we might get assurances for safe generalization behaviour, especially as AI systems become vastly more capable than humans.
Safety motivations aside, capable AI systems are extremely interesting objects of study, and doing digital neuroscience on them is comparatively much easier than studying biological neural systems.
A majority of reverse-engineering-focused interpretability work has involved sparse dictionary learning, which has various issues Sharkey et al., 2025. In response to these issues, my team and I at Goodfire (and previously Apollo Research) developed a new approach to decomposing neural networks, called Parameter Decomposition. We have used parameter decomposition to resolve feature splitting, identify attention-head distributed computations, identify circuits, and more.
We believe there are gaps in our method that remain to be solved. In our work (whether SAEs or parameter decomposition), we minimize 'description length'. But we are not yet confident that we have the right 'type of description'. We think understanding computational manifolds (which are projections of activation manifolds) is likely part of the answer here, since they may offer an even more concise description of neural computation than SDL latents or VPD parameter subcomponents.
MATS projects in my stream should primarily be aimed at improving methods for reverse engineering neural networks and at least be conceptually informed by parameter decomposition, manifolds, and minimum description length framings of interpretability, if not build on them directly.
Lee Sharkey is a Principal Investigator at Goodfire.
His team has focused on improved interpretability methods, including parameter decomposition methods such as Attribution-based Parameter Decomposition and Stochastic Parameter Decomposition and adVersarial Parameter Decomposition.
Previously, Lee was Chief Strategy Officer and cofounder of Apollo Research, and a Research Engineer at Conjecture, where he worked on sparse autoencoders as a solution to representational superposition.
Lucius Bushnaq is a Research Scientist at Goodfire.
He works on parameter decomposition methods for interpretability, such as Attribution-based Parameter Decomposition, Stochastic Parameter Decomposition, and adVersarial Parameter Decomposition. Alongside this, he works on learning theory and the theory behind interpretability, for example theoretical frameworks for computation in superposition and connections between singular learning theory and algorithmic information theory.
Previously, Lucius was a member of the interpretability team at Apollo Research, where he worked on the Local Interaction Basis and degeneracy in the loss landscape. He holds a PhD in physics.
Mentorship looks like a 1 h weekly meeting by default with approximately daily slack messages in between. Usually these meetings are just for updates about how the project is going, where I’ll provide some input and steering if necessary and desired. If there are urgent bottlenecks I’m more than happy to meet in between the weekly interval or respond on slack in (almost always) less than 24h. We'll often run daily standup meetings if timezones permit, but these are optional.
In general I'd like projects in my stream to at least be conceptually informed by parameter decomposition, manifolds, and minimum description length framings of interpretability, if not build on them directly.
Scholars and I will discuss projects and come to a consensus on what feels like a good direction. I will not tell scholars to work on a particular direction, since, in my experience, intrinsic motivation to work on a particular direction is important for producing good research.
The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.