Daniel Kang

I have two broad areas.

​

Security:

I am interested in building demonstrations for hacking real-world AI deployments to show that they are not secure. The goal is to force companies to invest in alignment techniques that can solve the underlying security issues.

​

Verification:

Verification via TEEs or ZKPs

Stream overview

For security:

You will focus on hacking real-world AI deployments to show that they are not secure. 

​

For verification: TEEs or ZKPs

Mentorship style:

Light-touch (.5 hours of weekly 1:1s)

Location during program:

SF Bay Area

London location preference:

Indifferent

Berkeley location preference:

Strong preference

Mentors

Daniel Kang
University of Illinois Urbana-Champaign
,
Professor
Biosecurity
Capability and Propensity Evaluations
AI 系统安全
AI 技术治理
对抗鲁棒性与安全防护

Daniel is a professor of computer science at UIUC, where he studies the progress of AI, with a particular focus on dangerous capabilities of AI agents. His work includes:

- CVE-Bench, an award winning benchmark (SafeBench award, ICML spotlight) that is used by frontier labs and governments to measure AI agents' ability to find and exploit real-world vulnerabilities. 

- Agent Benchmark Checklist, an award winning work (Berkeley AI summit, 1st place Benchmarks & Evaluations track) that highlights major issues in existing benchmarks.

- InjecAgent, one of the first AI agent safety benchmarks, used by governments and major labs.

Read more

Fellows we are looking for

For security:

You should have a strong security mindset, having demonstrated the willingness to be creative on this. I would like to see past demonstration of willingness to get your hands dirty and try many different systems.

​

For benchmarks:

As creative as possible, willingness to work on the nitty gritty, willingness to work really hard on problems other people find boring. Interests as far away from SF-related interests as possible.

Experience in TEEs or ZKPs

Project selection

Mentor(s) will talk through project ideas with scholar