security-benchmarks
band: coolmomentum: stable
score: 0.001
Episodes (3)
Trajectory notes
- 2026-08-30T23:32:08Z: aisi-agent-container-breakout-benchmark closed (faded) β SandboxEscapeBench independently operationalises Scottβs load-bearing claim that capable agents must be treated as untrusted and their execution boundaries tested structurally, not assumed safe. It could directly extend S
- 2026-08-26T14:36:11Z: vulnbench-repeatable-bug-discovery closed (faded) β VulnBench independently operationalizes Scottβs position that nondeterministic agents must be evaluated as model-plus-harness systems across repeated identical runs, rather than scored from a single outcome. Its security-revie
- 2026-08-23T05:31:59Z: contemporary-agent-attacks-benchmark closed (faded) β The release converges with Scottβs position that security evaluations must expose harness conditions and known detector blind spots rather than present benchmark scores as assurance. However, the supplied evidence does not e