vox-deorum's controlled CivBench claims GLM-5.3 now beats Opus 5.5 at long-horizon Civilization V play while Qwen-3.8-27B stays competitive; CivBench becoming a cited reference benchmark for long-horizon strategic planning across frontier and open-weight models โ or its GLM-over-Opus ranking failing replication โ resolves it.
state: watchingheat: mediumuncertainty: mediumnovelscott: highagent-evaluation open-weight-models long-horizon-agentsvox-deorum
What is this?
vox-deorum (a builder with a COLM 2026 paper) released a controlled version of CivBench โ a progress-based benchmark for LLM strategists playing multiplayer Civilization V (Vox Populi mod) via their Vox Deorum platform. The first-party results claim GLM-5.3 (open-weight) outperforms Opus 5.5 (frontier closed) on long-horizon strategic play, with Qwen-3.8-27B staying competitive. The benchmark measures model-under-agentic-setup, not raw foundation models, using turn-level progress signals because win/loss is too sparse over hundreds of turns. No independent replication exists yet; the striking open-beats-frontier ranking is falsifiable but unconfirmed.
Why it matters to Scott
A first-party long-horizon agent benchmark (CivBench) claiming an open-weight model (GLM-5.3) beats a frontier closed model (Opus 5.5) on multi-hundred-turn strategic play โ directly bearing on Scott's capability-symmetry and model-perishability theses, his model-plus-harness evaluation doctrine, and his long-running-agents architecture. The result is falsifiable but unreplicated; independent confirmation would constitute decision-useful evidence for open-weight competitiveness on long-horizon planning and for progress-based evaluation methodology. Scott's local-inference substrate (gamepc/Ollama) and trace-backed comparison harness make this concretely testable on his own stack.
ip:concept.capability-symmetryip:concept.model-plus-harness-benchmark-unitip:concept.evaluation-driven-developmentip:concept.capability-auditip:framework.long-running-agentsip:concept.measurable-convergenceip:concept.verification-loopsdev:concept.trace-backed-agent-comparisondev:concept.hardware-aware-local-inferencedev:project.gamepcradar:anthropic-opus55-cache-read-repricingradar:500-dollar-9b-rl-catalog-reviewradar:1password-scam-agent-benchmarkradar:agent-review-studio-local-evaluationradar:agentgauntlet-failure-benchmarkradar:ai-benchmark-saturation-distortionradar:activevision-repeated-perception-gapradar:aimee-native-model-memory
queries asked of Scott's wikis
- agent-evaluation benchmark methodology replication standards
- open-weight-models competitive parity frontier models long-horizon
- long-horizon-agents evaluation frameworks progress-based metrics
- local-inference economics open-model sovereignty
- agent-memory agent-maintained-wikis evaluation harnesses
- benchmark-harnesses civ-game-environments tool-mediated-agents
Measured heat
now 0 pts/hpeak 24 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 137h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p77 vs 1247 stories at the 96h mark (now 137h old) โ ahead of jenny-local-coding-harness (1.0x), behind claude-code-remote-attribution-injection (1.0x)
Evidence (2) โ โญ canonical anchor
Interpretation history
2026-10-08T04:49:30Z
grounded: novel/high โ A first-party long-horizon agent benchmark (CivBench) claiming an open-weight model (GLM-5.3) beats a frontier closed model (Opus 5.5) on multi-hundred-turn str
2026-10-08T04:36:09Z
Sustained multi-day velocity spikes and magnitude-valve eligibility show the CivBench launch is resonating beyond a single flash; the falsifiable open-beats-frontier claim (GLM-5.3 > Opus 5.5) on long-horizon planning by a COLM-credentialed builder now has enough sustained attention to warrant watching for independent replication or citation.
2026-10-05T23:38:18Z
case created โ First-party benchmark launch from a builder with a COLM 2026 lineage, carrying a falsifiable open-beats-frontier ranking on a hot evaluation topic despite low engagement.
Decision trace
- 10-08 17:55feedback_briefingScott vote via UI
- 10-08 17:07attention_communicatedvox-deorum's controlled CivBench (extending COLM 2026 work) reports GLM-5.3 outperforming Opus 5.5 on multi-hundred-turn Civilization V play, with Qwen-3.8-27B holding competitively. Reddit post
- 10-08 17:07attention_routeHigh relevance to Scott's capability-symmetry, model-perishability, and model-plus-harness theses; his local-inference substrate (gamepc/Ollama) and trace-backed comparison harness make this conc
- 10-08 15:50attention_routeHigh relevance to Scott's capability-symmetry, model-perishability, and model-plus-harness theses; his local-inference substrate (gamepc/Ollama) and trace-backed comparison harness make this conc
- 10-08 15:49attention_candidatecoverage review: magnitude valve eligible (multi-platform, top-decile engagement); not yet communicated
- 10-08 15:49repriceSustained multi-day velocity spikes and magnitude-valve eligibility show the CivBench launch is resonating beyond a single flash; the falsifiable open-beats-frontier claim (GLM-5.3 > Opus 5.5) on l
- 10-08 15:49groundA first-party long-horizon agent benchmark (CivBench) claiming an open-weight model (GLM-5.3) beats a frontier closed model (Opus 5.5) on multi-hundred-turn strategic play โ directly bearing on Scott&
- 10-07 10:23sensor_dirtycomment_update
- 10-07 06:24sensor_dirtyvelocity_spike
- 10-07 03:26sensor_dirtycomment_update
- 10-06 22:22sensor_dirtyvelocity_spike
- 10-06 18:20sensor_dirtycomment_update
- 10-06 14:21sensor_dirtyvelocity_spike
- 10-06 11:21sensor_dirtycomment_update
- 10-06 10:38createFirst-party benchmark launch from a builder with a COLM 2026 lineage, carrying a falsifiable open-beats-frontier ranking on a hot evaluation topic despite low engagement.