Shafi Goldwasser, Orr Paradise

This stream works on self-proving models: training a predictive model to prove that its own probabilistic claims are self-consistent, and building the verifier that checks them. 

It also studies verifier hacking, where training against a verifier with no soundness guarantee can select a model that games it, and builds harnesses for evaluating mathematical definitions.

Stream overview

Self Proving Model Training

&

Building Harnesses for Evaluating Math definitions

  1. Implement the probabilistic-consistency proof system. We recently constructed a proof system with which a predictive model can prove that its probabilistic claims are self-consistent, with the verifier explicitly designed for implementation. The goal of this project is to build that verifier, then train a small prover (order 10M-100M parameters) against it via Transcript Learning or RLVF, and measure verifiability, i.e., the fraction of outputs the verifier accepts. This connects to the aforementioned collaboration with Cambridge, which the scholar could join.
  2. Verifier hacking, in theory. RLVR as used in practice trains models against verifiers with no soundness guarantees. When does training against such a "verifier" result in a model that games it? In a stylized model, I have preliminary results showing that even benign training dynamics can inadvertently select a cheating prover, and that which prover is selected depends on fine features of the training setup (for instance, the noise structure). The direction is concrete but wide open: one could relax the fairly stringent assumptions the current theory asks for, or go empirical and test whether the phenomenon occurs in actual RLVR training. The destination is a characterization of when verifier-based training is sound: a theory of reward hacking for verifier-based training.
  3. Universal Self-Proving models. The theoretical guarantees of Self-Proving models currently hold when trained against one fixed verifier. Can a single model prove its answers to verifiers specified in-context, including ones unseen during training? Mostly empirical (train across a family of proof systems, test generalization to new ones), with theory questions available close-by: what does in-context generalization over verifiers require? Presumably, a restricted class of verifiers (e.g. low-depth circuits) but can we say something more?

The list above is a suggested menu, certainly not a mandate. I am a strong believer in projects driven by the student's curiosity. My preferred mode of advising, and the way my advisors have advised me, is to be of service to the student: to let them follow the directions they find most interesting, mysterious, or promising, and to be there as a sounding board, friendly critic, verifier, etc. That is, I will say something if a student is headed down a path I see no way out of, but I have been humbled before (a wonderful experience) by students whose intuition or other senses were sharper than mine, and who ended up with results that genuinely surprised me.

Mentors

Shafi Goldwasser (Shafi)
Simmons Resilience Research Pod
,
Professor of Electrical Engineering and Computer Science, UC Berkeley and MIT
Boston
No items found.

Shafi Goldwasser is the C. Lester Hogan Professor in Electrical Engineering and Computer Sciences at the University of California, Berkeley, the RSA Professor of Electrical Engineering and Computer Science at MIT, and directs the Simons Institute's resilience research pod. Goldwasser is a cryptographer, and a co-founder and chief scientist of Duality Technologies.

Read more
SF Bay Area
No items found.

Orr Paradise is a scientist at EPFL, appointed in its Theory of Machine Learning Laboratory and its Verification and Computer Architecture Lab. Paradise is also a researcher on Project CETI's theoretical analysis team, working on the machine analysis of sperm whale communication.

Read more

Mentorship style

Fellows we are looking for

Project selection

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

No items found.
SF Bay Area
Biosecurity
SF Bay Area
Founding and Field-Building
Systems Security
Systems Security
SF Bay Area
Empirical
SF Bay Area
Empirical
SF Bay Area
SF Bay Area
Empirical
SF Bay Area
SF Bay Area