ExploitGym is a benchmark designed to test whether AI agents can extend inputs that trigger known vulnerabilities into working exploits with concrete security impact. Led by Berkeley RDI at UC Berkeley with collaborators from the Max Planck Institute, UC Santa Barbara, Arizona State University, Anthropic, OpenAI, and Google, it contains 898 real-world tasks spanning userspace software, the V8 JavaScript engine, and the Linux kernel. The supplied snippets report frontier-model progress and at least one use in evaluating a Claude preview, but they primarily reflect the benchmark’s creators and do not establish independent validation, reproducibility, or measurement reliability; the alleged sandbox-escape incident appears only in a single secondary Substack snippet.
No intersection found in Scott’s wikis, and no radar pages currently track ExploitGym, CyberGym, or this validation question. The supplied evidence also does not yet establish the independent evaluations hypothesized by the case.
queries asked of Scott's wikis
- cyber-agent capability evaluations and benchmark validity
- agent harnesses for autonomous exploit generation
- sandboxing and containment for tool-using agents
- dual-use evaluations for frontier coding agents
- human guidance limits in autonomous cyber workflows
- reproducible benchmarks for long-horizon agent tasks
2026-07-29T12:35:00Z
No independent evaluation of ExploitGym has appeared despite repeated repricing; adjacent evidence only reinforces the gap between discovery and exploitation without testing the benchmark itself. The case has faded without progress.
2026-07-29T12:21:44Z
evidence attached: hn.story.49096301 — This is independent evidence that vulnerability discovery does not automatically translate into practical exploitation, a key question in the open case.
2026-07-29T12:21:44Z
evidence attached: hn.story.49096009 — Reports that AI-discovered vulnerabilities remain difficult to exploit materially support the case's distinction between finding bugs and turning them into working attacks.
2026-07-29T08:27:29Z
The source release makes ExploitGym inspectable and lowers the barrier to independent reproduction, but it is infrastructure for validation rather than validation itself. Until an unaffiliated evaluator publishes results, the measurement-reliability hypothesis remains untested.
2026-07-29T08:20:58Z
evidence attached: hn.story.49094493 — Directly provides the benchmark source needed for independent validation of ExploitGym.
2026-07-29T06:26:23Z
No newly identifiable independent evaluation, reproduction, or implementation tests ExploitGym itself; the apparent update only repeats the existing ambiguity and adjacent skepticism. The validation hypothesis remains untested and cold.
2026-07-28T21:28:05Z
No new independent evaluation, reproduction, or implementation is identifiable; the available discussion continues to repeat ambiguity and adjacent skepticism rather than test ExploitGym itself. The validation hypothesis remains untested and cold.
2026-07-28T20:27:51Z
The new reporting adds adjacent skepticism about converting AI-found vulnerabilities into working exploits, but it is not an independent evaluation of ExploitGym and does not directly validate or invalidate the benchmark. The core measurement-reliability hypothesis remains untested and cold.
2026-07-28T20:21:29Z
evidence attached: hn.story.49089211 — This reporting materially challenges the premise that AI-discovered vulnerabilities reliably translate into exploitable attacks.
2026-07-27T17:25:25Z
The attachment adds no identifiable independent evaluation, reproduction, or implementation, so it does not change the benchmark-validation question. Repeated discussion is only amplifying the same ambiguity around the reported OpenAI result.
2026-07-27T15:28:17Z
The newly attached material still provides no independent evaluation, reproduction, or implementation; it only repeats uncertainty about what the reported OpenAI result demonstrates. The benchmark-validation hypothesis remains untested and cold.
2026-07-27T12:26:45Z
The apparent update adds no independent evaluation, reproduction, or implementation; discussion still only questions what the reported model result demonstrates. This is repetitive amplification of the existing uncertainty, so the validation hypothesis remains untested.
2026-07-27T11:28:28Z
No independent evaluation, reproduction, or implementation has appeared; the attached material continues to amplify uncertainty about the reported result rather than validate ExploitGym’s measurement reliability.
2026-07-27T10:22:26Z
The added discussion exposes uncertainty about what the reported OpenAI result actually demonstrates, but supplies no independent evaluation, reproduction, or implementation evidence. The validation hypothesis therefore remains untested rather than corroborated.
2026-07-27T10:20:59Z
evidence attached: reddit.post.1v7va85 — The question directly bears on whether OpenAI models obtained ExploitGym solutions and how much autonomous exploitation the benchmark demonstrates.
2026-07-25T21:22:03Z
grounded: novel/none — No intersection found in Scott’s wikis, and no radar pages currently track ExploitGym, CyberGym, or this validation question. The supplied evidence also does no
2026-07-25T21:21:27Z
case created — This is a distinct end-to-end exploitation evaluation episode rather than another vulnerability-discovery claim, but it currently has only one low-engagement observation.