Caspar Oesterheld (Redwood, conceptual reasoning capabilities)

Theory of change: Soon, most important work will be done by AI. AI is going to increasingly advise people and help with important things, many of which are time-sensitive and path dependent, e.g., work on alignment/safety (including various things like how LLMs should behave given that they’re very persuasive); how to think about acausal trade; how to organize society. It seems good for AI to do well at those things.

Of course, a lot of the relevant skills for doing well at these tasks are the same skills that cause AI risk and that AI companies work on (and are incentivized to work on) by default; like coding, some kinds of forecasting, etc.

We want to make models better at things that are net positive for the future, but that likely won’t benefit much from said default training (or perhaps will even be made worse by such training – e.g., via sycophancy).

In practice, a lot of the tasks that we’re interested in from this perspective are what we call “conceptual”: tasks that are hard to verify and don't have clear ground truth but where we nonetheless feel like we can make progress through argument and reason.

You can visit conceptualreasoning.ai to get a sense of our work to date.

We also take a keen interest in projects directly aimed at making future acausal interactions go well.

Stream overview

We expect fellows to contribute to our research agenda of measuring and improving models' conceptual reasoning in domains that are neglected and especially important for making the future go well such as theoretical alignment, AI safety macrostrategy and decision theory. Fellows with the right skill set might work on projects that try to influence how AIs reason about decision theory and acausal interactions specifically although we expect only a small minority of potential fellows to fit this profile.

A large fraction of the work involves either a) creating datasets in a way that involves some conceptual reasoning on the side of the fellow or b) trying out various ideas for training and seeing if they improve performance on conceptual reasoning evals.

We expect to have new project ideas by the time the fellowship starts and fellows can also suggest their own projects. As illustrative examples, prospective applicants can check out our project ideas list from May:

https://docs.google.com/document/d/1iuYfsfPQNHIKGOttTDCpAf6xBHe6J3zyTqlRm973swQ/edit?tab=t.mznrcx1wk842#heading=h.dkc2x29shpk5

Mentors

Caspar Oesterheld
Redwood Research
,
Member of Technical Staff
SF Bay Area
Agent Foundations
AI Control and Monitoring
Structural Risk and Societal Dynamics

Caspar Oesterheld is a researcher at Redwood Research where he works on improving models' conceptual reasoning capabilities, i.e., their reasoning about questions where we cannot verify the answer and the best way to make progress is through argumentation. Much of this work has been done in collaboration with Anthropic.

Previously, he completed his computer science PhD at Carnegie Mellon University where he was assistant director of the Foundations of Cooperative AI Lab. Caspar has published about multi-agent AI interactions, decision theory in Newcomb-like decision problems, and, informally, about how models reason about conceptual questions. He has served as a research mentor for MATS, PIBBSS, CLR and astra (incoming).

Read more

Mentorship style

Fellows we are looking for

None

Project selection

Streams

The Winter 2026 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

SF Bay Area
Dangerous Capability Evals
Boston
Policy and Governance
Adversarial Robustness, Policy & Governance, Red-Teaming, Safeguards
New York City
Control, Scalable Oversight, Red-Teaming, Model Organisms, Monitoring
SF Bay Area
Policy and Governance
Policy & Governance
SF Bay Area
Control, Monitoring, Dangerous Capability Evals
SF Bay Area
Security, Compute Infrastructure
London
Theory
Interpretability
London
Scheming & Deception, Dangerous Capability Evals, Control, Red-Teaming
SF Bay Area
Dangerous Capability Evals, Red-Teaming, Model Organisms, Control, Monitoring
Toronto
Interpretability
London
Control, Monitoring, Safeguards, Dangerous Capability Evals, Scheming & Deception
Chicago
Biorisk, Security, Safeguards
SF Bay Area
Interpretability, Agent Foundations
London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting