2026-10-11 17:10 UTC

HunterBench claims its live-infrastructure benchmark can reproducibly compare LLM and agent pentesting capabilities beyond isolated vulnerability-exploitation tasks, giving security engineers a more realistic basis for model selection.

state: expiredheat: lowuncertainty: highknownscott: lowagentic-security pentesting-benchmarks cyber-agentsHunterBench

What is this?

HunterBench is presented as a benchmark for comparing LLMs and AI agents on penetration testing against live infrastructure, aiming to move beyond isolated vulnerability-exploitation tasks toward realistic, end-to-end attack workflows. The supplied results support the need for controlled, reproducible evaluations on realistic targets: existing agents perform much worse autonomously than with human guidance, especially on complex chained attacks. However, the snippets do not identify who operates HunterBench or provide direct evidence for its methodology, results, or reproducibility claims, so those details remain unverified here.

Why it matters to Scott

The core position is already explicit in Scott’s “Capability Audit,” “Benchmarking the Wrong Unit,” and “Model-Plus-Harness Benchmark Unit,” while the radar already tracks realistic agent evaluation under “agent-benchmarks” and offensive-security agents under “cyber-agents.” HunterBench is a domain-specific example rather than a demonstrated extension or challenge, and the supplied evidence does not verify its methodology, reproducibility, or results sufficiently to affect Scott’s model-selection or benchmark work.
ip:concept.capability-auditip:concept.benchmarking-the-wrong-unitip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisonradar:concept.agent-benchmarksradar:concept.cyber-agentsradar:concept.agent-evaluation
queries asked of Scott's wikis
  • live-environment benchmarks for autonomous agents
  • reproducible evaluation of stochastic agent systems
  • cyber-agent harnesses and human-in-the-loop autonomy
  • benchmarking end-to-end workflows versus isolated tasks
  • model selection through task-specific capability evaluations
  • agent observability and replay in stateful environments

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditWhich LLM is actually best at pentesting? benchmark to find out
LocalLLaMA
TomatoWasabi3926
🟧 echo.blog ⭐Introduces a benchmark intended to compare current LLMs and AI agents on pentesting against live infrastructure.HunterBench——

Interpretation history

Decision trace