The UK AI Security Institute claims its released benchmark can safely and reproducibly measure AI agents’ container-breakout capabilities, providing actionable evidence for sandbox evaluation and design.
state: expiredheat: lowuncertainty: highconvergesscott: highagentic-security agent-safety security-benchmarksUK AI Security Institute
What is this?
Researchers from the University of Oxford and the UK AI Security Institute released SandboxEscapeBench, an open benchmark for testing whether shell-enabled AI agents can escape Docker or Kubernetes containers and reach the host system. The supplied snippets describe 18 scenarios spanning orchestration, runtime, and kernel vulnerabilities, and report that advanced models frequently exploit common misconfigurations when prompted. To limit risk, the public release reportedly uses known vulnerability classes; the snippets support its use for sandbox evaluation but do not independently establish the stronger claim that it is fully safe or reproducible.
Why it matters to Scott
SandboxEscapeBench independently operationalises Scott’s load-bearing claim that capable agents must be treated as untrusted and their execution boundaries tested structurally, not assumed safe. It could directly extend SiloOS and Bubblewrap validation with adversarial container-escape fixtures and creates a dated-receipts opportunity, although the benchmark’s own safety and reproducibility claims still require independent verification.
ip:framework.siloosip:concept.runtime-containmentip:concept.model-plus-harness-benchmark-unitdev:project.silo-osdev:technology.bubblewrapdev:concept.trace-backed-agent-comparisonradar:concept.sandbox-escaperadar:concept.security-benchmarksradar:concept.benchmark-integrityradar:concept.agent-sandboxingradar:prime-intellect-offline-sandbox-escape
queries asked of Scott's wikis
- coding-agent sandbox threat models
- container isolation for autonomous agents
- agent harness privilege boundaries
- security benchmark validity and reproducibility
- sandbox escape evaluation integrity
- defense-in-depth for tool-using agents
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
| source | object | author | score | comments |
| 🟧 hn | A benchmark for safely measuring container breakout capabilities | paulpauper | 1 | 1 |
| 🟧 echo.paper ⭐ | The original research artifact is the arXiv v1 paper, “Quantifying Frontier LLM Capabilities for Container Sandbox Escape.” Its abstract say | Rahul Marchand, Art O Cathain, Jerome Wynne, Philippos Maximos Giavridis, Stuart Jennings, Freddy Tuxworth, Tolga H. Dur, Sam Deverett, John Wilkinson, Jason Gwartz, Harry Coppock | — | — |
Interpretation history
2026-08-30T23:32:07Z
Repeated checks have produced no independent replication, implementation report, or adoption, so the live episode has faded without validating the benchmark’s stronger safety and reproducibility claims. The artifact remains potentially useful and can be reopened if concrete external results emerge.
2026-08-28T23:24:09Z
No replication, implementation report, or independent assessment has appeared, so the benchmark remains a promising but unvalidated sandbox-testing artifact. This recheck adds only staleness and does not change its substantive meaning.
2026-08-26T22:42:50Z
The released benchmark is substantive enough to watch as an actionable sandbox-testing artifact, but this recheck adds no independent validation of its safety, reproducibility, or measurement quality. With no new evidence or discussion, the case cools while awaiting replication or implementation results.
2026-08-26T22:36:40Z
grounded: converges/high — SandboxEscapeBench independently operationalises Scott’s load-bearing claim that capable agents must be treated as untrusted and their execution boundaries test
2026-08-26T22:33:57Z
origin walked (codex/luna, conf 0.96): anchor hn.story.49456457 -> echo.paper.fe3f125bd6 by Rahul Marchand, Art O Cathain, Jerome Wynne, Philippos Maximos Giavridis, Stuart Jennings, Freddy Tuxworth, Tolga H. Dur, Sam Deverett, John Wilkinson, Jason Gwartz, Harry Coppock
2026-08-26T22:32:47Z
case created — A first-party benchmark for agent sandbox escape is a concrete security artifact with direct relevance to agent-execution infrastructure.
Decision trace
- 08-31 09:32expireRepeated checks have produced no independent replication, implementation report, or adoption, so the live episode has faded without validating the benchmark’s stronger safety and reproducibility claim
- 08-31 09:32alert_silentThe only delta is elapsed time, and the original release was already routed; there is no new consequential fact to interrupt Scott for.
- 08-31 09:32alert_routeThe only delta is elapsed time, and the original release was already routed; there is no new consequential fact to interrupt Scott for.
- 08-29 09:24repriceNo replication, implementation report, or independent assessment has appeared, so the benchmark remains a promising but unvalidated sandbox-testing artifact. This recheck adds only staleness and does
- 08-29 09:24alert_silentThere is no new consequential delta beyond elapsed time; the release was already routed, and another interruption should wait for independent validation or concrete adoption.
- 08-29 09:24alert_routeThere is no new consequential delta beyond elapsed time; the release was already routed, and another interruption should wait for independent validation or concrete adoption.
- 08-27 08:42repriceThe released benchmark is substantive enough to watch as an actionable sandbox-testing artifact, but this recheck adds no independent validation of its safety, reproducibility, or measurement quality.
- 08-27 08:42alert_silentThe benchmark release was already routed, and this delta contains only an unchanged reobservation; there is no new consequential fact that merits another interruption.
- 08-27 08:42alert_routeThe benchmark release was already routed, and this delta contains only an unchanged reobservation; there is no new consequential fact that merits another interruption.
- 08-27 08:37alert_shadowThe first-party release and research artifact establish a concrete, open benchmark spanning 18 container-escape scenarios across misconfiguration, privilege, kernel, runtime and orchestration weakness
- 08-27 08:37alert_routeThe first-party release and research artifact establish a concrete, open benchmark spanning 18 container-escape scenarios across misconfiguration, privilege, kernel, runtime and orchestration weakness
- 08-27 08:36groundSandboxEscapeBench independently operationalises Scott’s load-bearing claim that capable agents must be treated as untrusted and their execution boundaries tested structurally, not assumed safe. It co
- 08-27 08:33promote_anchororigin walk conf 0.96
- 08-27 08:32createA first-party benchmark for agent sandbox escape is a concrete security artifact with direct relevance to agent-execution infrastructure.