2026-10-11 17:11 UTC

CVE-Bench's publishers present a benchmark for evaluating AI agents' ability to exploit web vulnerabilities, potentially giving builders a task-specific measure of offensive agent capability.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagentic-security benchmarksCVE-Bench

What is this?

CVE-Bench is a sandbox benchmark for testing AI agents against 40 critical-severity vulnerabilities in real-world web applications, drawn from the National Vulnerability Database. Its code is hosted under uiuc-kang-lab; the paper lists Yuxuan Zhu, Daniel Kang, and collaborators, and Kang describes evaluations with vulnerability descriptions supplied (“one-day”) or withheld (“zero-day”). The supplied snippets also describe a public leaderboard and a v2.0 effort to prevent benchmark shortcuts, but do not establish the timing of the case’s announcement or provide overall agent success rates.

Why it matters to Scott

CVE-Bench gives Scott a concrete candidate testbed for extending the Security Reviewer Method’s requirement for reachable, independently checkable exploit paths into agent capability evaluation; its reported anti-shortcut work also converges with his Specification Gaming concern. This is distinct from the radar’s HunterBench and ExploitGym cases, but the supplied material establishes neither benchmark validity nor announcement timing, so it supports investigating an evaluation resource rather than claiming new capability gains or a dated-receipts publishing opportunity.
ip:source.security-reviewer-method-ebookip:concept.evaluation-driven-developmentip:concept.specification-gamingradar:hunterbench-live-pentesting-benchmarkradar:exploitgym-agent-exploitation-validationradar:concept.security-benchmarks
queries asked of Scott's wikis
  • agent evaluation harnesses real-world tasks success verification
  • reward hacking benchmark loopholes evaluation validity
  • coding agent sandboxing security permissions
  • autonomous agent exploration tool misuse failure analysis
  • offensive capability measurement dual-use penetration testing

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnCVE-Bench evaluates the capability of AI agents to exploit web vulnerabilitiessyumei20
🟧 echo.blog ⭐The project site is linked under the claim that CVE-Bench evaluates AI agents' capability to exploit web vulnerabilities.CVE-Bench——

Interpretation history

Decision trace