Truthful AI

This stream will focus on evaluating dangerous capabilities in language models and detecting deception and dishonesty.

Stream overview

  • Defining and evaluating situational awareness in LLMs (relevant paper)
  • Predicting the emergence of other dangerous capabilities in LLMs (e.g. deception, agency, misaligned goals)
  • Studying emergent reasoning at training time (“out-of-context” reasoning). See Reversal Curse.
  • Detecting deception and dishonesty in LLMs using black-box methods
  • Enhancing human epistemic abilities using LLMs (e.g., AutocastTruthfulQA)

Mentors

Owain Evans
Truthful AI
,
Research Lead
SF Bay Area
Misalignment Science
Capability and Propensity Evaluations

Owain has a broad interest in AI alignment and reducing AGI risk. He is investigating dangerous capabilities and the emergence of misalignment in LLMs, along with self-awareness and latent reasoning. Owain previously worked on AI deception (How to Catch an AI Liar), truthfulness (TruthfulQA), and the Reversal Curse. Owain runs an independent AI Safety non-profit, based at Constellation in Berkeley. He previously worked at the University of Oxford and at Ought. He has mentored 30+ junior AI Safety researchers through MATS and other programs.

Read more
Jan Betley
Truthful AI
,
Researcher
SF Bay Area
Misalignment Science

Jan worked as a software developer for over a decade before shifting to AI safety in 2023. He is an ARENA and Astra Fellowship alumni, interested in anything related to out-of-context reasoning in LLMs.

Read more

Mentorship style

Fellows we are looking for

Project selection

Streams

The Winter 2026 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

Oxford
Theory
AI Welfare
SF Bay Area
Control, Model Organisms, Scheming & Deception, Strategy & Forecasting
SF Bay Area
Theory
Interpretability
Tübingen
Dangerous Capability Evals, Agent Foundations, Adversarial Robustness, Monitoring, Scalable Oversight, Scheming & Deception
SF Bay Area
Policy and Governance
Dangerous Capability Evals, Policy & Governance
New York City
Monitoring, Dangerous Capability Evals, Scalable Oversight, Safeguards
SF Bay Area
Policy and Governance
Strategy & Forecasting, Policy & Governance
Montreal
Agent Foundations, Dangerous Capability Evals, Monitoring, Control, Red-Teaming, Scalable Oversight
SF Bay Area
Control, Model Organisms, Red-Teaming, Scheming & Deception