2026-10-11 17:12 UTC

The UK AI Security Institute claims its released benchmark can safely and reproducibly measure AI agents’ container-breakout capabilities, providing actionable evidence for sandbox evaluation and design.

state: expiredheat: lowuncertainty: highconvergesscott: highagentic-security agent-safety security-benchmarksUK AI Security Institute

What is this?

Researchers from the University of Oxford and the UK AI Security Institute released SandboxEscapeBench, an open benchmark for testing whether shell-enabled AI agents can escape Docker or Kubernetes containers and reach the host system. The supplied snippets describe 18 scenarios spanning orchestration, runtime, and kernel vulnerabilities, and report that advanced models frequently exploit common misconfigurations when prompted. To limit risk, the public release reportedly uses known vulnerability classes; the snippets support its use for sandbox evaluation but do not independently establish the stronger claim that it is fully safe or reproducible.

Why it matters to Scott

SandboxEscapeBench independently operationalises Scott’s load-bearing claim that capable agents must be treated as untrusted and their execution boundaries tested structurally, not assumed safe. It could directly extend SiloOS and Bubblewrap validation with adversarial container-escape fixtures and creates a dated-receipts opportunity, although the benchmark’s own safety and reproducibility claims still require independent verification.
ip:framework.siloosip:concept.runtime-containmentip:concept.model-plus-harness-benchmark-unitdev:project.silo-osdev:technology.bubblewrapdev:concept.trace-backed-agent-comparisonradar:concept.sandbox-escaperadar:concept.security-benchmarksradar:concept.benchmark-integrityradar:concept.agent-sandboxingradar:prime-intellect-offline-sandbox-escape
queries asked of Scott's wikis
  • coding-agent sandbox threat models
  • container isolation for autonomous agents
  • agent harness privilege boundaries
  • security benchmark validity and reproducibility
  • sandbox escape evaluation integrity
  • defense-in-depth for tool-using agents

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnA benchmark for safely measuring container breakout capabilitiespaulpauper11
🟧 echo.paper ⭐The original research artifact is the arXiv v1 paper, “Quantifying Frontier LLM Capabilities for Container Sandbox Escape.” Its abstract sayRahul Marchand, Art O Cathain, Jerome Wynne, Philippos Maximos Giavridis, Stuart Jennings, Freddy Tuxworth, Tolga H. Dur, Sam Deverett, John Wilkinson, Jason Gwartz, Harry Coppock——

Interpretation history

Decision trace