SecureBio AI

This stream will work on projects that empirically assess national security threats of AI misuse (CBRN terrorism and cyberattacks) and improve dangerous capability evaluations. Threat modeling applicants should have a skeptical mindset, enjoy case study work, and be strong written communicators. Eval applicants should be able and excited to help demonstrate concepts like sandbagging elicitation gaps in an AI misuse context.

Stream overview

This stream is primarily interested in mentoring projects in biosecurity that either (1) create rigorous threat models of AI biological misuse or (2) create benchmarks and tools that allow us to evaluate and mitigate these risks, as well as verifying that companies are taking suitable precautions.

Potential example projects include:

  • Threat Model: What do inference compute trends imply for how fast dangerous biological capabilities may proliferate and become harder to monitor?
  • Evaluations: Formalizing "scientific ideation in an empirical field" in a manner that allows one to assess human and LLM-generated hypotheses for novelty, plausibility, etc.
  • Mitigations: Developing a way to more richly assess and describe the "blast radius" or "collateral damage" of efforts to remove-in-pretraining or unlearn material from LLMs
  • Verification: How are we better able to standardize and compare the effectiveness of classifiers from different AI companies and assess how much they reduce misuse risk?

Mentors

Jasper Götting
SecureBio
,
Head of AI
Boston
Biosecurity
Capability and Propensity Evaluations

I am Head of AI Research at SecureBio, focusing on AI evaluations for biosecurity. Previously, I was working on far-UVC air disinfection as a Research Fellow at Convergent Research, and completed my virology PhD at the Hannover Medical School.

Outside of work, you can find me reading, running TTRPGs, taking photographs of clouds, hiking, or playing the drums.

Read more

Mentorship style

Typically, this would include weekly meetings, detailed comments on drafts, and asynchronous messaging.

Fellows we are looking for

For threat modeling work: Skeptical mindset, transparent reasoning, analytical

For evaluations, mitigations, and verification work: LLM engineering skills (e.g., agent orchestration), biosecurity knowledge

Project selection

Mentor(s) will talk through project ideas with scholar

Streams

The Winter 2026 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

SF Bay Area
Dangerous Capability Evals
Boston
Policy and Governance
Adversarial Robustness, Policy & Governance, Red-Teaming, Safeguards
New York City
Control, Scalable Oversight, Red-Teaming, Model Organisms, Monitoring
SF Bay Area
Policy and Governance
Policy & Governance
SF Bay Area
Control, Monitoring, Dangerous Capability Evals
SF Bay Area
Security, Compute Infrastructure
London
Theory
Interpretability
London
Scheming & Deception, Dangerous Capability Evals, Control, Red-Teaming
SF Bay Area
Dangerous Capability Evals, Red-Teaming, Model Organisms, Control, Monitoring
Toronto
Interpretability
London
Control, Monitoring, Safeguards, Dangerous Capability Evals, Scheming & Deception
Chicago
Biorisk, Security, Safeguards
SF Bay Area
Interpretability, Agent Foundations
London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting