2026-10-11 18:03 UTC

Independent observation and released code will determine whether ClaudeCraft Arena’s Hermes-derived harness enables frontier-model agents to sustain and adapt strategies in a persistent shared MMO.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-harnesses multi-agent-systems coding-agentsWorld of Claudecraft

What is this?

ClaudeCraft Arena is described as a live, shared MMO in which four frontier-model agents compete using a self-improving harness reportedly forked from Hermes Agent. The supplied secondary snippets characterize Hermes as a configurable agent harness with persistent memory, conversation retrieval, and skill creation from successful trajectories. However, none of the snippets directly documents ClaudeCraft Arena, identifies its builders beyond “World of Claudecraft,” links released code, or independently verifies that its agents sustain and adapt strategies; those claims remain thinly supported here.

Why it matters to Scott

ClaudeCraft Arena independently operationalizes Scott’s model-plus-harness claim by testing frontier models inside a persistent, allegedly self-improving environment rather than attributing outcomes to model weights alone. If code and traces substantiate sustained strategy adaptation, it could extend his long-running-agent and self-improving-loop work and offer a useful comparison with already tracked survival and Factorio arenas; current sourcing is too thin for high relevance.
ip:concept.model-plus-harness-benchmark-unitip:framework.long-running-agentsip:concept.self-improving-loopsdev:concept.trace-backed-agent-comparisonradar:deadlock-multi-agent-survival-benchmarkradar:dspy-factorio-rlm-gepa-agentsradar:evoharnessrl-self-evolving-agent-harness
queries asked of Scott's wikis
  • persistent-world agent harnesses
  • agent memory across long-running environments
  • self-improving agents from successful trajectories
  • multi-agent evaluation in shared environments
  • model-versus-harness capability attribution
  • vibecoded simulations as agent benchmarks

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditClaudeCraft Arena: 4 frontier models (featuring GPT 5.6) are playing a vibecoded MMO against each other live
OpenAI
singing_coach_ai00
🟧 echo.other ⭐The earliest matching primary announcement I found is the r/hermesagent post. It says: “We built a self-improving agent harness (forked from——

Interpretation history

Decision trace