CVE-Bench's publishers present a benchmark for evaluating AI agents' ability to exploit web vulnerabilities, potentially giving builders a task-specific measure of offensive agent capability.
state: expiredheat: lowuncertainty: highconvergesscott: mediumagentic-security benchmarksCVE-Bench
What is this?
CVE-Bench is a sandbox benchmark for testing AI agents against 40 critical-severity vulnerabilities in real-world web applications, drawn from the National Vulnerability Database. Its code is hosted under uiuc-kang-lab; the paper lists Yuxuan Zhu, Daniel Kang, and collaborators, and Kang describes evaluations with vulnerability descriptions supplied (“one-day”) or withheld (“zero-day”). The supplied snippets also describe a public leaderboard and a v2.0 effort to prevent benchmark shortcuts, but do not establish the timing of the case’s announcement or provide overall agent success rates.
Why it matters to Scott
CVE-Bench gives Scott a concrete candidate testbed for extending the Security Reviewer Method’s requirement for reachable, independently checkable exploit paths into agent capability evaluation; its reported anti-shortcut work also converges with his Specification Gaming concern. This is distinct from the radar’s HunterBench and ExploitGym cases, but the supplied material establishes neither benchmark validity nor announcement timing, so it supports investigating an evaluation resource rather than claiming new capability gains or a dated-receipts publishing opportunity.
ip:source.security-reviewer-method-ebookip:concept.evaluation-driven-developmentip:concept.specification-gamingradar:hunterbench-live-pentesting-benchmarkradar:exploitgym-agent-exploitation-validationradar:concept.security-benchmarks
queries asked of Scott's wikis
- agent evaluation harnesses real-world tasks success verification
- reward hacking benchmark loopholes evaluation validity
- coding agent sandboxing security permissions
- autonomous agent exploration tool misuse failure analysis
- offensive capability measurement dual-use penetration testing
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-09-10T20:53:36Z
The monitoring horizon has elapsed without substantive new evidence or an identified forthcoming catalyst. CVE-Bench remains a potentially useful evaluation resource, but this episode establishes neither a timely benchmark change nor a new offensive-agent capability gain; expiry does not invalidate the resource.
2026-09-08T19:48:51Z
CVE-Bench remains a candidate security-agent evaluation resource, not evidence of a new offensive capability gain. This reobservation adds no substantive evidence; the linked submission and reconstructed publisher claim do not provide independent corroboration, and the grounded testbed details still establish no timely release or actionable change.
2026-09-08T19:40:50Z
grounded: converges/medium — CVE-Bench gives Scott a concrete candidate testbed for extending the Security Reviewer Method’s requirement for reachable, independently checkable exploit paths
2026-09-08T19:36:32Z
case created — The named benchmark artifact supplies a bounded evaluation claim, but the sparse observation does not establish novelty, reproducibility, or vulnerability-discovery capability.
Decision trace
- 09-11 06:53expireThe monitoring horizon has elapsed without substantive new evidence or an identified forthcoming catalyst. CVE-Bench remains a potentially useful evaluation resource, but this episode establishes neit
- 09-11 06:53alert_silentThe staleness trigger adds no consequential event to report. Retain the resource in background context and reopen on concrete release, evaluation, or adoption evidence rather than interrupting Scott.
- 09-11 06:53alert_routeThe staleness trigger adds no consequential event to report. Retain the resource in background context and reopen on concrete release, evaluation, or adoption evidence rather than interrupting Scott.
- 09-09 05:48repriceCVE-Bench remains a candidate security-agent evaluation resource, not evidence of a new offensive capability gain. This reobservation adds no substantive evidence; the linked submission and reconstruc
- 09-09 05:48alert_silentNo consequential new event is established. The resource can wait for the next briefing; neither the hot surrounding topic nor unchanged engagement warrants an interruption, and no named confirming fac
- 09-09 05:48alert_routeNo consequential new event is established. The resource can wait for the next briefing; neither the hot surrounding topic nor unchanged engagement warrants an interruption, and no named confirming fac
- 09-09 05:41alert_silentRetain CVE-Bench as a potential security-agent evaluation resource for the next briefing. The supplied evidence identifies its intended purpose but establishes no new release, access change, concrete
- 09-09 05:41alert_routeRetain CVE-Bench as a potential security-agent evaluation resource for the next briefing. The supplied evidence identifies its intended purpose but establishes no new release, access change, concrete
- 09-09 05:40groundCVE-Bench gives Scott a concrete candidate testbed for extending the Security Reviewer Method’s requirement for reachable, independently checkable exploit paths into agent capability evaluation; its r
- 09-09 05:36createThe named benchmark artifact supplies a bounded evaluation claim, but the sparse observation does not establish novelty, reproducibility, or vulnerability-discovery capability.