Independent runs will determine whether Deadlock’s unrestricted 12-agent survival arena reveals reproducible coordination and emergent-strategy failures that conventional task benchmarks miss.
state: expiredheat: lowuncertainty: highknownscott: lowagent-benchmarks multi-agent-systems agent-harnesses
What is this?
Deadlock appears to be an “AI Squid Game” arena in which 12 autonomous agents compete in survival games, each operating independently, with elimination framed as death. The supplied snippets do not identify its creators or provide results from independent runs, so the claim that it reproducibly exposes coordination or emergent-strategy failures is not yet established. A separate Stanford benchmark, Agent Island, supports the broader premise that dynamic multiplayer environments can test cooperation, conflict, and persuasion beyond static benchmarks, but it does not validate Deadlock specifically.
Why it matters to Scott
Scott already holds the relevant position: non-deterministic multi-agent behavior requires repeated, observable runs rather than outcome-only demonstrations, as captured in “Non-Determinism,” “Agent Observability,” and his LLM self-play work. The radar also already tracks nearly identical replication questions in “Open-Ended Agent Coordination Benchmark” and “Vending-Bench 2 Agent Collusion”; without independent results, traces, or a distinctive evaluation method, Deadlock is another candidate arena rather than a consequential update.
ip:concept.non-determinismip:concept.agent-observabilitydev:concept.llm-self-play-refinementradar:open-ended-agent-coordination-benchmarkradar:vending-bench-2-agent-collusion
queries asked of Scott's wikis
- dynamic multi-agent benchmarks vs static task benchmarks
- reproducibility of emergent agent behavior
- agent harnesses for adversarial multi-agent simulations
- measuring coordination and strategy failures
- benchmark contamination and saturation resistance
- observability and replay for multi-agent runs
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (3) — ⭐ canonical anchor
Interpretation history
2026-08-06T12:24:50Z
No independent runs, traces, or implementations emerged within the observation window, leaving Deadlock as an unvalidated creator demo rather than a developing benchmark episode. The broader benchmark idea remains relevant, but this case has not earned further attention.
2026-08-02T19:21:30Z
No independent run, trace, or implementation has appeared; the apparent update is only reobservation of the same creator-originated material. Deadlock remains an unvalidated arena concept rather than evidence of reproducible multi-agent failure modes.
2026-08-02T18:22:11Z
The new post clarifies the arena’s implementation but is another creator-originated description, not an independent run or observable trace. Deadlock remains a candidate benchmark whose claimed value depends on external replication and evidence of reproducible failure modes.
2026-08-02T18:21:21Z
evidence attached: reddit.post.1vdp28y — Directly describes the Deadlock benchmark's unrestricted 12-agent survival arena and provides useful implementation context.
2026-08-02T17:21:29Z
The added material remains creator-originated and supplies no independent runs, traces, or distinctive evaluation method, so it does not advance the reproducibility hypothesis. With engagement unchanged and no corroboration, this is a candidate arena awaiting implementation evidence rather than a moving benchmark episode.
2026-08-02T16:26:19Z
grounded: known/low — Scott already holds the relevant position: non-deterministic multi-agent behavior requires repeated, observable runs rather than outcome-only demonstrations, as
2026-08-02T16:24:01Z
origin walked (codex/luna, conf 0.98): anchor reddit.post.1vdm33u -> echo.youtube.f44743a280 by ronydkid
2026-08-02T16:22:05Z
case created — Deadlock is a distinct, technically substantive benchmark episode, though evidence is currently limited to its creator’s low-engagement introduction.
Decision trace
- 08-06 22:24expireNo independent runs, traces, or implementations emerged within the observation window, leaving Deadlock as an unvalidated creator demo rather than a developing benchmark episode. The broader benchmark
- 08-03 05:21repriceNo independent run, trace, or implementation has appeared; the apparent update is only reobservation of the same creator-originated material. Deadlock remains an unvalidated arena concept rather than
- 08-03 05:20mark_dirtyengagement_update
- 08-03 04:22repriceThe new post clarifies the arena’s implementation but is another creator-originated description, not an independent run or observable trace. Deadlock remains a candidate benchmark whose claimed value
- 08-03 04:21attachDirectly describes the Deadlock benchmark's unrestricted 12-agent survival arena and provides useful implementation context.
- 08-03 04:20propose_attachDirectly describes the Deadlock benchmark's unrestricted 12-agent survival arena and provides useful implementation context.
- 08-03 03:21repriceThe added material remains creator-originated and supplies no independent runs, traces, or distinctive evaluation method, so it does not advance the reproducibility hypothesis. With engagement unchang
- 08-03 03:20mark_dirtyengagement_update
- 08-03 02:26groundScott already holds the relevant position: non-deterministic multi-agent behavior requires repeated, observable runs rather than outcome-only demonstrations, as captured in “Non-Determinism,” “Agent O
- 08-03 02:24promote_anchororigin walk conf 0.98
- 08-03 02:22createDeadlock is a distinct, technically substantive benchmark episode, though evidence is currently limited to its creator’s low-engagement introduction.