UK AISI Control

This stream is for the UK AISI Red-team. The team focuses on stress-testing mitigations for AI risk, including misuse safeguards, control techniques and model alignment red-teaming. We plan to work on projects building and improving methods for performing these kinds of evaluations and methods.

Stream overview

Project could include: 

  •  developing methods for automated red-teaming of mitigations for any of the risk areas
  •  building evaluation environments for evaluating misuse of agentic AI systems
  •  developing attacks against asynchronous or stateful monitoring systems for misuse
  •  developing and red-teaming control protocols and modelling effect in realistic deployment situations
  •  improving automated auditing tools such as Petri and Bloom to be more realistic, controllable, and with additional beneficial affordances for propensity evaluations; building propensity evaluations within automated auditing tools or otherwise."

Mentors

Asa Cooper Stickland
UK AISI
,
Research Scientist
London
Misalignment Science
AI Control and Monitoring
Capability and Propensity Evaluations
Adversarial Robustness and Safeguards

Asa Strickland is a research scientist at the UK AI Security Institute working on AI control red teaming and model organisms of misalignment. He was previously a postdoc with Sam Bowman at NYU, did MATS with Owain Evans, and mentored for the MATS, SPAR and Pivotal fellowships. He got his PhD at the University of Edinburgh, supervised by Iain Murray.

Read more

Mentorship style

Each scholar will have one primary mentor from the Red Team who will provide weekly guidance and day-to-day support

Scholars will also have access to secondary advisors within their specific sub-team (misuse, alignment, or control) for technical deep-dives

Team lead Xander Davies and advisors Geoffrey Irving and Yarin Gal will provide periodic feedback through team meetings and project reviews

For scholars working on cross-cutting projects, we can arrange mentorship from multiple sub-teams as needed

Structure: 

Weekly 1:1 meetings (60 minutes) with primary mentor for project updates, technical guidance, and problem-solving

Asynchronous communication via Slack/email throughout the week for quick questions and feedback

Bi-weekly team meetings where scholars can present work-in-progress and get broader team input

Working style:

We expect scholars to work semi-independently – taking initiative on their research direction while leveraging mentors for guidance on technical challenges, research strategy, and navigating AISI resources

Scholars will have access to our compute resources and operational support to focus on research

We encourage scholars to document their work and, if appropriate, aim for publication or public blog posts

Fellows we are looking for

We're looking for scholars with hands-on experience in machine learning and AI security, particularly those interested in adversarial robustness, red teaming, or AI safeguards. Ideal candidates would have:

  • Experience with large language models (training, fine-tuning, evaluation, or safety research)
  • Strong technical foundations in ML, ideally with coding experience in PyTorch or Inspect
  • Interest in one or more of our three focus areas: misuse (securing systems against bad actors), alignment (ensuring AI systems behave as intended), or control (keeping AI systems under human control even when misaligned)
  • A mission-driven mindset and curiosity about how AI security research can inform real-world policy and deployment decisions
  • An ability to advocate for your own research ideas and work in a self-directed way, while also collaborating effectively and prioritizing team efforts over extensive solo work.

We welcome scholars at various career stages especially those who are eager to work on problems with direct impact on how frontier AI is governed and deployed.

Project selection

 Scholars will choose from a set of predefined project directions aligned with our current research priorities, such as:

  • Developing automated methods to test AI misuse safeguards
  • Investigating data poisoning attacks and defenses
  • Designing benchmarks for misuse detection across multiple model interactions
  • Testing control measures for potentially misaligned AI systems

 We'll provide initial direction and guidance on project scoping, then scholars will have autonomy to explore specific approaches within that framework.

 Expect weekly touchpoints to ensure progress and refine directions. 

 If mentees have particular ideas they're excited about that they see as fitting within the scope of the team's work, they're welcome to propose them, but there is no guarantee they will be selected

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

Grand Rapids
Agent Foundations
Washington, D.C.
Compute Infrastructure, Policy & Governance, Security
London
Dangerous Capability Evals, Compute Infrastructure, Policy & Governance, Strategy & Forecasting
Washington, D.C.
Compute Infrastructure, Security
London
Control, Monitoring
London
Control, Scheming & Deception, Dangerous Capability Evals, Model Organisms, Monitoring
SF Bay Area
Security, Dangerous Capability Evals
SF Bay Area
Dangerous Capability Evals
Boston
Adversarial Robustness, Policy & Governance, Red-Teaming, Safeguards
New York City
Control, Scalable Oversight, Red-Teaming, Model Organisms, Monitoring
SF Bay Area
Policy & Governance
SF Bay Area
Control, Monitoring, Dangerous Capability Evals
SF Bay Area
Security, Compute Infrastructure
London
Interpretability
London
Scheming & Deception, Dangerous Capability Evals, Control, Red-Teaming
SF Bay Area
Dangerous Capability Evals, Red-Teaming, Model Organisms, Control, Monitoring
Toronto
Interpretability
London
Control, Monitoring, Safeguards, Dangerous Capability Evals, Scheming & Deception
Chicago
Biorisk, Security, Safeguards
SF Bay Area
Interpretability, Agent Foundations