Independent repeated-run evaluations will determine whether VulnBench reproducibly measures how consistently LLM security agents rediscover the same vulnerabilities.
state: expiredheat: lowuncertainty: highconvergesscott: mediumagentic-security coding-agents security-benchmarksVulnBench
What is this?
Snyk’s VulnBench JS 1.0 evaluates the repeatability of agentic LLM security reviews by running 300 vulnerability-finding scans with the same code, prompt, and harness. Snyk reports that reference-matched findings were relatively stable, while additional model-generated findings varied widely between runs; its best configuration achieved 75.4% F1 against Snyk’s deterministic SAST reference. The supplied results provide related evidence that reliable vulnerability detection remains difficult, but they do not establish an independent reproduction of VulnBench itself or identify the arXiv paper’s authors.
Why it matters to Scott
VulnBench independently operationalizes Scott’s position that nondeterministic agents must be evaluated as model-plus-harness systems across repeated identical runs, rather than scored from a single outcome. Its security-review setting also bears directly on his trace-backed agent comparisons and bounded WordPress security-review workflow, but the supplied evidence is still Snyk’s own report and provides no independent reproduction, limiting its present weight.
ip:concept.model-plus-harness-benchmark-unitip:concept.non-determinismip:concept.evaluation-driven-developmentdev:concept.trace-backed-agent-comparisondev:project.wordpress-security-reviewradar:dfah-bench-agent-trajectory-driftradar:visa-agentic-sast-harness-validationradar:concept.agent-evaluationradar:concept.agent-reliabilityradar:concept.security-agents
queries asked of Scott's wikis
- stochastic agent evaluation and repeated-run reliability
- coding-agent benchmark design and reproducible harnesses
- security agents versus deterministic static analysis
- variance-aware scoring for nondeterministic AI systems
- LLM vulnerability discovery and false-positive triage
- agent reliability across identical runs
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-26T14:36:11Z
After repeated checks, VulnBench still has no independent reproduction, implementation uptake, or substantive methodological critique. The benchmark remains potentially useful but has not developed into an active evaluation standard, so this episode can fade until genuinely new evidence appears.
2026-08-24T14:28:40Z
No independent reproduction, implementation, or technical critique has emerged; repeated unchanged observations leave VulnBench as a first-party benchmark artifact rather than evidence of a reproducible evaluation standard.
2026-08-22T13:36:01Z
The case remains a first-party benchmark claim awaiting independent repeated-run validation; the latest reobservation adds no substantive corroboration or implementation evidence.
2026-08-20T12:45:44Z
No independent repeated-run reproduction has appeared; the case remains a first-party benchmark awaiting external validation, with no new evidence beyond the original artifact.
2026-08-20T12:37:50Z
grounded: converges/medium — VulnBench independently operationalizes Scott’s position that nondeterministic agents must be evaluated as model-plus-harness systems across repeated identical
2026-08-20T12:35:31Z
origin walked (codex/luna, conf 0.98): anchor hn.story.49373152 -> echo.paper.3c5f3d8ace by Liran Tal, Johannes Kloos, Arsenii Rudich, Stephen Thoemmes, and Manoj Nair
2026-08-20T12:33:42Z
case created — The dedicated benchmark is a concrete evaluation artifact focused on the undermeasured repeatability of agentic vulnerability discovery.
Decision trace
- 08-27 00:36expireAfter repeated checks, VulnBench still has no independent reproduction, implementation uptake, or substantive methodological critique. The benchmark remains potentially useful but has not developed in
- 08-27 00:36alert_silentThe latest change is only negligible engagement on an otherwise unchanged artifact; independent replication, adoption, or a material technical critique can reopen the case through routine monitoring.
- 08-27 00:36alert_routeThe latest change is only negligible engagement on an otherwise unchanged artifact; independent replication, adoption, or a material technical critique can reopen the case through routine monitoring.
- 08-25 00:28repriceNo independent reproduction, implementation, or technical critique has emerged; repeated unchanged observations leave VulnBench as a first-party benchmark artifact rather than evidence of a reproducib
- 08-25 00:28alert_silentThe staleness trigger adds no consequential evidence. An independent repeated-run reproduction, benchmark adoption, or material methodological critique can surface through routine review.
- 08-25 00:28alert_routeThe staleness trigger adds no consequential evidence. An independent repeated-run reproduction, benchmark adoption, or material methodological critique can surface through routine review.
- 08-22 23:36repriceThe case remains a first-party benchmark claim awaiting independent repeated-run validation; the latest reobservation adds no substantive corroboration or implementation evidence.
- 08-22 23:36alert_silentOnly negligible engagement changed, with no independent reproduction or new technical result; this can wait for routine review.
- 08-22 23:36alert_routeOnly negligible engagement changed, with no independent reproduction or new technical result; this can wait for routine review.
- 08-20 22:45repriceNo independent repeated-run reproduction has appeared; the case remains a first-party benchmark awaiting external validation, with no new evidence beyond the original artifact.
- 08-20 22:45alert_silentThe reobservation is unchanged and adds no consequential delta; independent reproduction or a materially new implementation can surface through normal review.
- 08-20 22:45alert_routeThe reobservation is unchanged and adds no consequential delta; independent reproduction or a materially new implementation can surface through normal review.
- 08-20 22:43alert_silentThe paper provides a substantive repeated-run security-agent evaluation directly relevant to Scott’s benchmark methodology, but it was published weeks ago and presents no immediate access, security, p
- 08-20 22:43surface_candidateThe paper provides a substantive repeated-run security-agent evaluation directly relevant to Scott’s benchmark methodology, but it was published weeks ago and presents no immediate access, security, p
- 08-20 22:43alert_routeThe paper provides a substantive repeated-run security-agent evaluation directly relevant to Scott’s benchmark methodology, but it was published weeks ago and presents no immediate access, security, p
- 08-20 22:37groundVulnBench independently operationalizes Scott’s position that nondeterministic agents must be evaluated as model-plus-harness systems across repeated identical runs, rather than scored from a single o
- 08-20 22:35promote_anchororigin walk conf 0.98
- 08-20 22:33createThe dedicated benchmark is a concrete evaluation artifact focused on the undermeasured repeatability of agentic vulnerability discovery.