Independent observation and released code will determine whether ClaudeCraft Arena’s Hermes-derived harness enables frontier-model agents to sustain and adapt strategies in a persistent shared MMO.
state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-harnesses multi-agent-systems coding-agentsWorld of Claudecraft
What is this?
ClaudeCraft Arena is described as a live, shared MMO in which four frontier-model agents compete using a self-improving harness reportedly forked from Hermes Agent. The supplied secondary snippets characterize Hermes as a configurable agent harness with persistent memory, conversation retrieval, and skill creation from successful trajectories. However, none of the snippets directly documents ClaudeCraft Arena, identifies its builders beyond “World of Claudecraft,” links released code, or independently verifies that its agents sustain and adapt strategies; those claims remain thinly supported here.
Why it matters to Scott
ClaudeCraft Arena independently operationalizes Scott’s model-plus-harness claim by testing frontier models inside a persistent, allegedly self-improving environment rather than attributing outcomes to model weights alone. If code and traces substantiate sustained strategy adaptation, it could extend his long-running-agent and self-improving-loop work and offer a useful comparison with already tracked survival and Factorio arenas; current sourcing is too thin for high relevance.
ip:concept.model-plus-harness-benchmark-unitip:framework.long-running-agentsip:concept.self-improving-loopsdev:concept.trace-backed-agent-comparisonradar:deadlock-multi-agent-survival-benchmarkradar:dspy-factorio-rlm-gepa-agentsradar:evoharnessrl-self-evolving-agent-harness
queries asked of Scott's wikis
- persistent-world agent harnesses
- agent memory across long-running environments
- self-improving agents from successful trajectories
- multi-agent evaluation in shared environments
- model-versus-harness capability attribution
- vibecoded simulations as agent benchmarks
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-13T21:33:33Z
The launch claim has produced no code, traces, independent observation, or reported agent behavior within the initial observation window. Without a concrete follow-up path, this remains an unvalidated experiment rather than a developing signal and can be reopened if artifacts emerge.
2026-08-11T20:44:33Z
No new code, traces, independent observation, or implementation evidence has appeared; the case remains a thinly sourced live experiment whose sustained-strategy and self-improvement claims are unvalidated.
2026-08-11T20:32:58Z
grounded: converges/medium — ClaudeCraft Arena independently operationalizes Scott’s model-plus-harness claim by testing frontier models inside a persistent, allegedly self-improving enviro
2026-08-11T20:27:32Z
origin walked (codex/luna, conf 0.83): anchor reddit.post.1vls7kj -> echo.other.c228fc3219 by None
2026-08-11T20:25:21Z
case created — The reported live, open-source multi-model environment is a concrete agent-runtime artifact, but currently has only one low-engagement source and no linked repository.
Decision trace
- 08-14 07:33expireThe launch claim has produced no code, traces, independent observation, or reported agent behavior within the initial observation window. Without a concrete follow-up path, this remains an unvalidated
- 08-14 07:33alert_silentThere is no new consequential delta; max staleness merely confirms that the thin launch evidence has not developed, so another briefing would add no value.
- 08-14 07:33alert_routeThere is no new consequential delta; max staleness merely confirms that the thin launch evidence has not developed, so another briefing would add no value.
- 08-12 06:44repriceNo new code, traces, independent observation, or implementation evidence has appeared; the case remains a thinly sourced live experiment whose sustained-strategy and self-improvement claims are unvali
- 08-12 06:44alert_silentThis look adds no consequential delta beyond the already-routed launch announcement; unchanged engagement does not justify another alert.
- 08-12 06:44alert_routeThis look adds no consequential delta beyond the already-routed launch announcement; unchanged engagement does not justify another alert.
- 08-12 06:36alert_shadowThe builders have announced and opened a live, observable multi-agent experiment using a Hermes-derived harness, directly relevant to Scott’s work on persistent agents and model-versus-harness attribu
- 08-12 06:36alert_routeThe builders have announced and opened a live, observable multi-agent experiment using a Hermes-derived harness, directly relevant to Scott’s work on persistent agents and model-versus-harness attribu
- 08-12 06:32groundClaudeCraft Arena independently operationalizes Scott’s model-plus-harness claim by testing frontier models inside a persistent, allegedly self-improving environment rather than attributing outcomes t
- 08-12 06:27promote_anchororigin walk conf 0.83
- 08-12 06:25createThe reported live, open-source multi-model environment is a concrete agent-runtime artifact, but currently has only one low-engagement source and no linked repository.