Anthropic
EXPLOITBENCH Capability Ladder Benchmark for LLM Cybersecurity Agents
Pages
28
Time to read
84 mins
Publication
Language
English
Pages
28
Time to read
84 mins
Publication
Language
English
This document is a preprint presenting EXPLOITBENCH, a benchmark designed to evaluate the capabilities of large language models (LLMs) in constructing exploits against production software. It addresses the limitations of existing benchmarks that treat exploitation as a binary event, instead proposing a graded approach that decomposes exploitation into 16 measurable flags. These flags range from basic bug coverage to advanced capabilities like arbitrary code execution. The authors detail their methodology, which includes using a deterministic oracle for verification and testing against known vulnerabilities in the V8 JavaScript engine. The results indicate a significant capability divide between publicly available LLMs and private models, with only one public model achieving arbitrary code execution under specific conditions. The study concludes that while public models can trigger crashes, they struggle to build the necessary primitives for advanced exploitation, highlighting the emerging challenges in this area of cybersecurity.