The hardest problems in AI safety may not be solvable with experiments alone. Instead, they may require the kind of foundational thinking in mathematics and philosophy that gives the field something solid to build on. Streams in this track work on agent foundations, formal models of trust and agency, mechanistic interpretability theory, and AI welfare. We're looking for researchers with deep mathematical maturity who want to tackle the problems that will still matter when AI systems are far more capable than they are today.
This track works on problems where the goal is durable conceptual progress rather than experimental results on today's models. The assumption is that some of the hardest alignment issues, such as questions around agency, optimization, trust, and the structure of cognition, will not be settled by simply scaling current empirical techniques, and that mathematical and philosophical foundations will matter significantly when AI systems are much more capable than they are now. Projects here cover agent foundations, formal models of trust and agency, mechanistic interpretability theory, and AI welfare. Methods rely mainly on reasoning and proof, though some work intersects with empirical interpretability or formal verification.
We are looking for fellows with serious mathematical maturity and a willingness to sit with problems where the right formalization is itself part of the work. Essential traits are research independence (theory questions are open-ended and require self-direction), fluency with formal reasoning (proofs, probability, type theory, dynamical systems, or analogous), and the ability to write clearly about abstract ideas. Strong candidates have come from mathematics, theoretical computer science, theoretical physics, formal philosophy, and economic theory, but a fellow’s background is not as significant as their demonstrated ability to do hard formal work, which can ideally be documented in written output we can read.
Fellows are matched to mentors based on fit and produce concrete artifacts (e.g., papers, technical reports, conceptual write-ups, or formal results) by the end of the program . Target audiences for the work produced in this track include the agent foundations and alignment theory communities, alignment-relevant teams at frontier labs, and academic venues for formal work. Theory outputs typically have longer time horizons than empirical ones, and we expect many fellows to continue refining results past the conclusion of the program.
Agent Foundations research focused on clarifying conditions under which humans can justifiably trust artificial intelligence systems. When should one boundedly rational learning-theoretic process come to trust another?
We can discuss this more and decide on a different structure, but by default, 1 hour 1-on-1 meetings with each scholar once a week, plus a 2 hour group meeting which may also include outside collaborators.
Quality of fit is roughly proportional to philosophical skill times mathematical skill. Someone with excellent philosophical depth and almost no mathematics could be an OK fit, but would probably struggle to produce or evaluate proofs. Someone with excellent mathematical depth but no philosophy could be an OK fit, but might struggle to understand what assumptions and theorems are useful/interesting.
There will be some flexibility about what specific projects scholars will pursue. Abram will discuss the current state of his research with scholars and what topics scholars are interested in, aiming to settle on a topic by or before week 2.
The Alignment Research Center is a small non-profit research group based in Berkeley, California, that is working on a systematic and theoretically grounded approach to mechanistically explaining neural network behavior. We are interested in fellows with a strong math background and mathematical maturity. If you'd be excited to work on the research direction described in this blog post – then we'd encourage you to apply!
Scholars will work out of ARC's offices in Berkeley. Each scholar will meet with their mentor at least once a week for an hour, though 2-3 hours per week is not uncommon. Besides time with their official mentor, scholars will likely spend time working in collaboration with other researchers; a typical scholar will likely spend about 25% of their time actively collaborating or learning about others' research.
Each scholar will be paired with the mentor that best suits their skills and interests. The mentor will discuss potential projects with the scholar, and they will decide what project makes the most sense, based on ARC's research goals and the scholar's preferences.
Most scholars will work on multiple projects over the course of their time at ARC, and some scholars will work with multiple mentors.
Theory of change: Soon, most important work will be done by AI. AI is going to increasingly advise people and help with important things, many of which are time-sensitive and path dependent, e.g., work on alignment/safety (including various things like how LLMs should behave given that they’re very persuasive); how to think about acausal trade; how to organize society. It seems good for AI to do well at those things.
Of course, a lot of the relevant skills for doing well at these tasks are the same skills that cause AI risk and that AI companies work on (and are incentivized to work on) by default; like coding, some kinds of forecasting, etc.
We want to make models better at things that are net positive for the future, but that likely won’t benefit much from said default training (or perhaps will even be made worse by such training – e.g., via sycophancy).
In practice, a lot of the tasks that we’re interested in from this perspective are what we call “conceptual”: tasks that are hard to verify and don't have clear ground truth but where we nonetheless feel like we can make progress through argument and reason.
You can visit conceptualreasoning.ai to get a sense of our work to date.
We also take a keen interest in projects directly aimed at making future acausal interactions go well.
None
Projects in this stream will be on AI welfare and moral status; more specifically, on what it takes to be a moral patient and how we can determine whether AI systems meet the conditions. I'm looking for applicants who have ideas about these topics and are motivated to explore them in more detail.
By default, scholars will meet with me online for 1hr/week and I will respond to questions on email/slack.
I will talk through project ideas with scholar
Existing frameworks for understanding intelligent agency don't do a great job at describing multi-agent dynamics (e.g. agents recursively modeling each other, merging with each other, threatening each other, etc). Most work in my stream aims (implicitly or explicitly) to move towards a multi-agent understanding of intelligence.
I'll come meet scholars in person around 2 days a week on average. On those days I'll be broadly available for discussions and brainstorming. On other days scholars can message me for guidance (though I'd prefer to spend most of my effort on this during the in-person days).
No required qualifications—I'm looking for scholars who are capable of very clear and curious thinking, but I'm open to many ways that they might demonstrate that.
I will talk through project ideas with the scholar.
MATS 项目是一项为期 10 周的研究奖学金计划,旨在培养和支持从事人工智能对齐、透明度和安全领域工作的新兴研究人员。研究员将与世界一流的导师合作,获得专门的研究管理支持,并加入位于伯克利、致力于推动人工智能安全与可靠发展的活跃社区。该项目提供开展高影响力研究并开启人工智能安全领域长期职业生涯所需的架构、资源和指导。
MATS 导师均为来自人工智能安全、对齐、治理、领域建设及安全等广泛领域的顶尖研究人员。他们包括学术界人士、行业研究员以及独立专家,负责指导学者开展研究项目、提供反馈,并助力每位学者的研究成长。导师们的专业领域涵盖:
查看 往届及现任导师
关键日期
申请:
主项目将于 9 月 28 日至 12 月 4 日进行,获选研究员的延展阶段将于 12 月开始。
MATS 欢迎来自不同学术和专业背景的申请者——从机器学习、数学和计算机科学,到政策、经济学、物理学、认知科学、生物学和公共卫生,同时也欢迎没有传统研究背景的创业者、运营人员和领域建设者。主要要求是具备为人工智能安全做出贡献的强烈动机,并展现出技术能力、研究潜力或相关的运营经验。具备人工智能安全相关经验会有所帮助,但并非必要条件。