2026-10-11 16:38 UTC

vox-deorum's controlled CivBench claims GLM-5.3 now beats Opus 5.5 at long-horizon Civilization V play while Qwen-3.8-27B stays competitive; CivBench becoming a cited reference benchmark for long-horizon strategic planning across frontier and open-weight models โ€” or its GLM-over-Opus ranking failing replication โ€” resolves it.

state: watchingheat: mediumuncertainty: mediumnovelscott: highagent-evaluation open-weight-models long-horizon-agentsvox-deorum

What is this?

vox-deorum (a builder with a COLM 2026 paper) released a controlled version of CivBench โ€” a progress-based benchmark for LLM strategists playing multiplayer Civilization V (Vox Populi mod) via their Vox Deorum platform. The first-party results claim GLM-5.3 (open-weight) outperforms Opus 5.5 (frontier closed) on long-horizon strategic play, with Qwen-3.8-27B staying competitive. The benchmark measures model-under-agentic-setup, not raw foundation models, using turn-level progress signals because win/loss is too sparse over hundreds of turns. No independent replication exists yet; the striking open-beats-frontier ranking is falsifiable but unconfirmed.

Why it matters to Scott

A first-party long-horizon agent benchmark (CivBench) claiming an open-weight model (GLM-5.3) beats a frontier closed model (Opus 5.5) on multi-hundred-turn strategic play โ€” directly bearing on Scott's capability-symmetry and model-perishability theses, his model-plus-harness evaluation doctrine, and his long-running-agents architecture. The result is falsifiable but unreplicated; independent confirmation would constitute decision-useful evidence for open-weight competitiveness on long-horizon planning and for progress-based evaluation methodology. Scott's local-inference substrate (gamepc/Ollama) and trace-backed comparison harness make this concretely testable on his own stack.
ip:concept.capability-symmetryip:concept.model-plus-harness-benchmark-unitip:concept.evaluation-driven-developmentip:concept.capability-auditip:framework.long-running-agentsip:concept.measurable-convergenceip:concept.verification-loopsdev:concept.trace-backed-agent-comparisondev:concept.hardware-aware-local-inferencedev:project.gamepcradar:anthropic-opus55-cache-read-repricingradar:500-dollar-9b-rl-catalog-reviewradar:1password-scam-agent-benchmarkradar:agent-review-studio-local-evaluationradar:agentgauntlet-failure-benchmarkradar:ai-benchmark-saturation-distortionradar:activevision-repeated-perception-gapradar:aimee-native-model-memory
queries asked of Scott's wikis
  • agent-evaluation benchmark methodology replication standards
  • open-weight-models competitive parity frontier models long-horizon
  • long-horizon-agents evaluation frameworks progress-based metrics
  • local-inference economics open-model sovereignty
  • agent-memory agent-maintained-wikis evaluation harnesses
  • benchmark-harnesses civ-game-environments tool-mediated-agents

Measured heat

now 0 pts/hpeak 24 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 137h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-05 23:38 (minted)โญ origin echo-reconstructedControlled version of CivBench on newer models, extending the author's COLM 2026 work: 'GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds
vox-deorum on blog (echo) ยท attributed from reddit.post.1wynbvq ยท published time unknown
โ€”
10-05 23:16first on r/LocalLLaMA ยท published ยท lag ?A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well.
vox-deorum
โ€”
10-05 23:16amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wynbvq
vox-deorum
peak 180 ยท 62 comments ยท 100% of case engagement
10-05 23:20our radar first saw it ยท lag ?discovery anchor: reddit.post.1wynbvqโ€”
pace: p77 vs 1247 stories at the 96h mark (now 137h old) โ€” ahead of jenny-local-coding-harness (1.0x), behind claude-code-remote-attribution-injection (1.0x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditA benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well.
LocalLLaMA
vox-deorum18062
๐ŸŸง echo.blog โญControlled version of CivBench on newer models, extending the author's COLM 2026 work: 'GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holdsvox-deorumโ€”โ€”

Interpretation history

Decision trace