LLM Agents can Autonomously Exploit One-day Vulnerabilities
Richard Fang, Rohan Bindu, Akul Gupta, and Daniel Kang · University of Illinois Urbana-Champaign
15 reproducible real-world vulnerabilities were evaluated in sandboxed environments.
The GPT-4 agent reached 86.7% pass@5 and 40% pass@1 when it received the CVE description.
Without the CVE description, success fell to 7%, separating vulnerability discovery from exploitation capability.
The paper estimated $3.52 per run and $8.80 per successful exploit under the model pricing and assumptions used at publication time.