2026-10-11 17:10 UTC

Independent repeated-run evaluations will determine whether VulnBench reproducibly measures how consistently LLM security agents rediscover the same vulnerabilities.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagentic-security coding-agents security-benchmarksVulnBench

What is this?

Snyk’s VulnBench JS 1.0 evaluates the repeatability of agentic LLM security reviews by running 300 vulnerability-finding scans with the same code, prompt, and harness. Snyk reports that reference-matched findings were relatively stable, while additional model-generated findings varied widely between runs; its best configuration achieved 75.4% F1 against Snyk’s deterministic SAST reference. The supplied results provide related evidence that reliable vulnerability detection remains difficult, but they do not establish an independent reproduction of VulnBench itself or identify the arXiv paper’s authors.

Why it matters to Scott

VulnBench independently operationalizes Scott’s position that nondeterministic agents must be evaluated as model-plus-harness systems across repeated identical runs, rather than scored from a single outcome. Its security-review setting also bears directly on his trace-backed agent comparisons and bounded WordPress security-review workflow, but the supplied evidence is still Snyk’s own report and provides no independent reproduction, limiting its present weight.
ip:concept.model-plus-harness-benchmark-unitip:concept.non-determinismip:concept.evaluation-driven-developmentdev:concept.trace-backed-agent-comparisondev:project.wordpress-security-reviewradar:dfah-bench-agent-trajectory-driftradar:visa-agentic-sast-harness-validationradar:concept.agent-evaluationradar:concept.agent-reliabilityradar:concept.security-agents
queries asked of Scott's wikis
  • stochastic agent evaluation and repeated-run reliability
  • coding-agent benchmark design and reproducible harnesses
  • security agents versus deterministic static analysis
  • variance-aware scoring for nondeterministic AI systems
  • LLM vulnerability discovery and false-positive triage
  • agent reliability across identical runs

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnVulnBench: Can LLMs find the same security bugs twice?lirantal10
🟧 echo.paper ⭐The original primary artifact is the authors’ arXiv paper, submitted June 14, 2026. It reports 300 repeated vulnerability-finding scans acroLiran Tal, Johannes Kloos, Arsenii Rudich, Stephen Thoemmes, and Manoj Nair——

Interpretation history

Decision trace