I work on the science of evaluating advanced AI systems for biological and CBRN risks, with a particular interest in translating technical evidence into decisions by governments and frontier AI developers. In this stream, I’m interested in developing novel capability evaluations, studying how dangerous or dual-use capabilities diffuse into increasingly accessible models, and building scalable red-teaming methods that produce rigorous, decision-relevant evidence without requiring risky real-world demonstrations.
I expect to refine project selection with fellows at the beginning of the program based on their technical backgrounds, emerging research opportunities, and which questions seem most likely to produce decision-relevant results. Example projects could include:
Sunishchal Dev is an AI safety researcher working pre deployment testing with a focus on biosecurity at the U.S. AI Safety Institute within NIST’s Center for AI Standards and Innovation (CAISI). Previously, he was an AI Evaluations Research Scientist at RAND, where he led machine learning engineering efforts focused on evaluating frontier AI systems, including building biological capability benchmarks, assessing risks from open-weight models, and developing methods to make LLM-based evaluations more reliable. He was a fellow during MATS 6.0 under the mentorship of Marius Hobbhahn. Before moving into AI safety, Dev spent nearly a decade as a data scientist, machine learning engineer, and management consultant.
The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.