Sydney Von Arx

I work on adversarial stress-testing at METR. Fellows in this stream will build model organisms of misalignment that require real cognition to decide whether to misbehave, then test which detection methods actually catch them, with the goal of turning whatever works into an alignment evaluation we can run on real models.

Stream overview

I work on adversarial stress-testing at METR. The projects I am most excited about right now:

  • Detecting misaligned cognition in model organisms. We will have a suite of model organisms that require substantial cognition to determine whether or not to misbehave, alongside matched controls. The work is to apply existing detection methods and variants of them (reading chain-of-thought, out-of-distribution detection with probes, off-policy probes for things like "reasoning about the input", and fuzzing) and find out what actually catches these models. A hard constraint is that the auditor is never allowed to trigger the misbehavior: every input it uses must avoid the trigger, and it must still detect the cognition. The goal is to build an alignment evaluation on top of what works and apply it to real models.
  • Benchmarking the capability to scheme undetected. Building a benchmark for the level of reasoning a model would have to be capable of in order to scheme without getting caught.

Mentors

Sydney Von Arx
Nightingale
,
Technical Staff
SF Bay Area
Misalignment Science
AI Control and Monitoring
Capability and Propensity Evaluations

Sydney works on adversarial stress-testing at METR. She studied computational biology at Stanford and co-founded the Atlas Fellowship.

Read more

Mentorship style

Fellows we are looking for

Project selection

Streams

The Winter 2026 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

SF Bay Area
Dangerous Capability Evals
Boston
Policy and Governance
Adversarial Robustness, Policy & Governance, Red-Teaming, Safeguards
New York City
Control, Scalable Oversight, Red-Teaming, Model Organisms, Monitoring
SF Bay Area
Policy and Governance
Policy & Governance
SF Bay Area
Control, Monitoring, Dangerous Capability Evals
SF Bay Area
Security, Compute Infrastructure
London
Theory
Interpretability
London
Scheming & Deception, Dangerous Capability Evals, Control, Red-Teaming
SF Bay Area
Dangerous Capability Evals, Red-Teaming, Model Organisms, Control, Monitoring
Toronto
Interpretability
London
Control, Monitoring, Safeguards, Dangerous Capability Evals, Scheming & Deception
Chicago
Biorisk, Security, Safeguards
SF Bay Area
Interpretability, Agent Foundations
London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting