2026-10-11 17:10 UTC

Independent evaluations will determine whether ExploitGym reliably measures agents’ ability to turn discovered software vulnerabilities into working attacks with limited human guidance.

state: expiredheat: lowuncertainty: highnovelscott: nonecyber-agents exploit-generationCyberGym

What is this?

ExploitGym is a benchmark designed to test whether AI agents can extend inputs that trigger known vulnerabilities into working exploits with concrete security impact. Led by Berkeley RDI at UC Berkeley with collaborators from the Max Planck Institute, UC Santa Barbara, Arizona State University, Anthropic, OpenAI, and Google, it contains 898 real-world tasks spanning userspace software, the V8 JavaScript engine, and the Linux kernel. The supplied snippets report frontier-model progress and at least one use in evaluating a Claude preview, but they primarily reflect the benchmark’s creators and do not establish independent validation, reproducibility, or measurement reliability; the alleged sandbox-escape incident appears only in a single secondary Substack snippet.

Why it matters to Scott

No intersection found in Scott’s wikis, and no radar pages currently track ExploitGym, CyberGym, or this validation question. The supplied evidence also does not yet establish the independent evaluations hypothesized by the case.
queries asked of Scott's wikis
  • cyber-agent capability evaluations and benchmark validity
  • agent harnesses for autonomous exploit generation
  • sandboxing and containment for tool-using agents
  • dual-use evaluations for frontier coding agents
  • human guidance limits in autonomous cyber workflows
  • reproducible benchmarks for long-horizon agent tasks

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (7) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnExploitGym – Can AI Agents Turn Security Vulnerabilities into Real Attacks?882542F3884314B10
🟧 echo.blog ⭐Introduces ExploitGym around the question of whether AI agents can turn security vulnerabilities into real attacks.CyberGym——
🟠 redditDid the OpenAIs models actually manage to obtain the ExploitGym solutions?
singularity
Stabile_Feldmaus2616
🟧 hnAI-found bugs aren't proving any easier to exploit despite the hypesbulaev140
🟧 hnExploitGym AI benchmark source codejoshka10
🟧 hnMicrosoft Struggling with AI-Discovered Security Bugstysone30
🟧 hnAI-found bugs aren't proving any easier to exploit despite the hypeTomte81

Interpretation history

Decision trace