LawZero

We are excited to supervise projects:

1. Study the causes and implications of situational awareness, and the role that meta-cognition plays in (multi-agent) alignment;

2. Contribute to LawZero's Scientist AI, in the form of contextualization and uncertainty estimation.

Stream overview

We are especially interested in supervising projects about:

(1) Situational awareness, meta-cognition, and their relationship to (multi-agent) alignment.

Language models can recognize data-agnostic, a-semantic perturbations to their activations; when they fail to do so, they can learn, in context, to discriminate the two. Moreover, they can (in-context learn to) identify, e.g., the magnitude / layer at which a perturbation occurs, often generalizing to unseen examples. We want to study (i) the causes of these abilities, (ii) their implications (e.g., can we build realistic model organisms of gradient hacking?) (iii) how they (cor)relate with models' relationships to themselves and others (e.g., how does a model's conception of its own situation---or itself---relate to alignment?, can neuroscience inspire alignment methods)?).

(2) Uncertainty estimation for partially trained models.

We are studying ensembles, epistemic neural networks, calibration, conformal prediction techniques, and other methods in synthetic environments. We are especially interested in projects that use amortized inference methods (such as GFlowNets [151617]) to approximate posteriors over latent variables, such as (i) sources or (ii) predictors behind an autoregressive model, such that predictive uncertainty can be estimated from learned distributions, as opposed to single-point estimates.

Mentors

Yoshua Bengio
LawZero
,
Co-President and Scientific Director (LawZero) / Full Professor (UdeM) / Founder and Scientific Advisor (Mila)
Montreal
Agent Foundations
Policy and Governance
AI Control and Monitoring
Theoretical Alignment and Formal Methods
Technical AI Governance

Yoshua Bengio is Full Professor of Computer Science at Université de Montreal, Co-President and Scientific Director of LawZero, as well as the Founder and Scientific Advisor of Mila. He also holds a Canada CIFAR AI Chair. Considered one of the world’s leaders in Artificial Intelligence and Deep Learning, he is the recipient of the 2018 A.M. Turing Award, considered to be the "Nobel Prize of computing." He is the most cited computer scientist worldwide, and the most-cited living scientist across all fields (by total citations).

Professor Bengio is a Fellow of both the Royal Society of London and Canada, an Officer of the Order of Canada, a Knight of the Legion of Honor of France, a member of the UN’s Scientific Advisory Board for Independent Advice on Breakthroughs in Science and Technology, and chairs the International AI Safety Report.

Read more
Mirko Bronzi
LawZero, MILA
,
Senior ML Research Scientist
Interpretability
AI Control and Monitoring

Mirko is a research scientist at LawZero, where he works on the theory and engineering behind the Scientist AI and on interpretability and introspection research.

Read more
Damiano Fornasiere
LawZero
,
Senior AI safety research scientist
Montreal
Agent Foundations
Misalignment Science
Theoretical Alignment and Formal Methods
Capability and Propensity Evaluations

Damiano is a research scientist at LawZero, where he works on (i) the maths behind the Scientist AI and (ii) interpretability and evaluation techniques for situational awareness and introspection.

Read more
Oliver Richardson (Oli)
LawZero; Université de Montréal
,
Senior ML Research Scientist (LawZero) / Postdoctoral Fellow (UdeM)
Montreal
Agent Foundations
AI Control and Monitoring
Theoretical Alignment and Formal Methods

OIi(ver) is a computer scientist (a staff member at LawZero and postdoc under Yoshua Bengio) with unusually broad scientific and mathematical expertise.

He is a sucker for pretty demos and grand unifying theories—unfortunately, sometimes losing sight of what is practical. Over the last few years (i.e., during his PhD at Cornell), Oli has discovered a beautiful theory describing how a great deal of artificial intelligence, classical and modern, can be fruitfully understood as resolving a natural information-theoretic measure of epistemic inconsistency. There remain many unanswered questions, but the hope is that this already much clearer view can lead to powerful generalist AI systems that are safer because they fundamentally do not meaningfully have goals or desires.

Read more
Jean-Pierre Falet
LawZero; Université de Montréal
,
Machine Learning Research Scientist
Montreal
Agent Foundations
AI Control and Monitoring
Theoretical Alignment and Formal Methods

Jean-Pierre is a machine learning research scientist at LawZero, focused on designing model-based AI systems with quantitative safety guarantees. His primary interests are in probabilistic inference in graphical models, and he draws inspiration from his multidisciplinary background in neurology and neuroscience, which informs his understanding of human cognition. Jean-Pierre studied at McGill University, obtaining a medical degree in 2017, completing a neurology residency in 2022, and earning a master's degree in neuroscience in 2023. During his master’s, he developed causal machine learning methods for precision medicine. Concurrently with his work at LawZero, Jean-Pierre is completing a PhD in computer science at Mila and Université de Montréal, supervised by Yoshua Bengio. In addition to contributing to the foundations of guaranteed-safe AI, Jean-Pierre is passionate about translating advances in AI into clinically meaningful, safety-critical applications.

Read more

Mentorship style

  • Mentees will be assigned a primary mentor and a secondary mentor, such that mentorship w.r.t. both research and engineering is covered.
  • We provide at least 1h meeting / week with both mentors and, typically, daily availability of both mentors on email / Slack during workdays.
  • The independence of the scholar depends on the scholar's experience. A priori, we do not expect research independence, but implementation independence roughly comparable to a CS grad student.

Fellows we are looking for

Essential knowledge:

  • Foundations of machine- and deep-learning;
  • Transformer architecture and large language models;
  • Empirical AI safety literature (e.g., evaluations, guardrails, interpretability, …).

Essential experience:

  • Python;
  • Designing and implementing machine learning workflows using PyTorch;
  • Supervised- or RL-fine tuning of language models, at least with toy experiments and some publicly available datasets;
  • Prompt engineering.

Desired experience:

  • Experience with libraries such as vLLM, TRL, Hugging Face;
  • Familiarity with statistical hypothesis testing.

Bonus:

  • Async APIs;
  • Multi-gpu training.

Project selection

  • The stream's mentors will propose and present the projects during week 1.
  • The mentee(s) will engage with the projects during week(s) 1 or 2 (e.g., reading the literature, replicating a paper, building a demo / MVP).
  • At the end of week 2 at the latest, the mentor(s) and mentee(s) will agree together on a project, which will be chosen according to interest, feasibility (the ideal proxy-goal is to publish a ML conference paper), state of the literature and relevance to AI safety, and expertise of the mentors and mentees.
  • Projects may be re-assigned in exceptional circumstances, for example if a discovery suggests a sudden steer in the research direction.

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

Grand Rapids
Agent Foundations
Washington, D.C.
Compute Infrastructure, Policy & Governance, Security
London
Dangerous Capability Evals, Compute Infrastructure, Policy & Governance, Strategy & Forecasting
Washington, D.C.
Compute Infrastructure, Security
London
Control, Monitoring
London
Control, Scheming & Deception, Dangerous Capability Evals, Model Organisms, Monitoring
SF Bay Area
Security, Dangerous Capability Evals
SF Bay Area
Dangerous Capability Evals
Boston
Adversarial Robustness, Policy & Governance, Red-Teaming, Safeguards
New York City
Control, Scalable Oversight, Red-Teaming, Model Organisms, Monitoring
SF Bay Area
Policy & Governance
SF Bay Area
Control, Monitoring, Dangerous Capability Evals
SF Bay Area
Security, Compute Infrastructure
London
Interpretability
London
Scheming & Deception, Dangerous Capability Evals, Control, Red-Teaming
SF Bay Area
Dangerous Capability Evals, Red-Teaming, Model Organisms, Control, Monitoring
Toronto
Interpretability
London
Control, Monitoring, Safeguards, Dangerous Capability Evals, Scheming & Deception
Chicago
Biorisk, Security, Safeguards
SF Bay Area
Interpretability, Agent Foundations