Izolight's released Render Arena โ blind A/B voting over 1,200+ agent Blender-modeling runs comparing harnesses (pi, opencode, codex, Claude Code, dsh) and script-writing versus MCP integration โ becomes a used public benchmark for isolating harness-versus-model effects in agentic tooling; sustained external votes and builder citations confirm it, a stalled solo site closes it.
state: seedheat: lowuncertainty: mediumconvergesscott: highagent-evaluation agent-harnesses creative-tool-agentsIzolight
What is this?
Render Arena (render-arena.izolight.xyz) is a live, open benchmark attributed to Izolight where visitors blind A/B-vote on agent-produced Blender modeling results across the agent stack โ the site lists pi, omp, omp-mcp, omp-mcp-skill, codex, and opencode, with votes bound to anonymous browser tokens, rate-limited per address, and runs reproducible locally through containerized submissions that record the full Blender environment. The omp/omp-mcp/omp-mcp-skill variant naming is consistent with an integration-style (script vs MCP) comparison axis, but the case's claimed 1,200+ run count, the 'Claude Code' and 'dsh' entries, and any external traction are not established by these snippets โ no third-party coverage of Render Arena itself appears. What the search does surface is a parallel cluster of harness-isolating evaluation efforts in the coding domain: Ondemand-OSS's Harness Arena (LMSYS-style blind judging with an Elo leaderboard, model held constant), d1-m4ss/harness-benchmark, Supabase Evals, and Terminal-Bench โ so harness-vs-model isolation is emerging as a benchmark category, while Render Arena's own uptake remains unmeasured.
Why it matters to Scott
Izolight independently built the public instrument for Scott's model-plus-harness position โ blind A/B voting with the model held constant across the exact harnesses Scott runs (pi, opencode, codex, Claude Code) โ and the omp vs omp-mcp arms are a live, votable test of his code-execution-beats-MCP claim in a creative 3D domain his own fixed-fixture comparisons never covered. Traction is minimal today, but if the arena draws use it becomes dated receipts for his arguments or a public challenge to them, whichever way the votes fall.
ip:concept.model-plus-harness-benchmark-unitip:source.why-code-execution-beats-mcpip:source.give-the-agent-a-workshop-ebookdev:concept.trace-backed-agent-comparisondev:concept.brief-ab-testingradar:concept.agent-harnessesradar:concept.agent-evaluationradar:ship-harness-benchradar:swe-bench-pro-harness-cost-parityradar:frontierharness-17x-cost-variationradar:willison-astra-blender-workflow
queries asked of Scott's wikis
- harness versus model isolating eval agent benchmark
- MCP vs script tool integration agent design
- blind voting arena Elo eval methodology
- Blender agent creative tool automation
- pi coding agent harness notes
- containerized reproducible agent benchmark runs
Measured heat
now 0 pts/hpeak 2 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 139h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p56 vs 1247 stories at the 96h mark (now 139h old) โ ahead of ai-agent-ransomware-operation (1.0x), behind anthropic-fourth-cyber-incident-review-miss (1.0x)
Evidence (1) โ โญ canonical anchor
Interpretation history
2026-10-06T01:06:41Z
grounded: converges/high โ Izolight independently built the public instrument for Scott's model-plus-harness position โ blind A/B voting with the model held constant across the exact harn
2026-10-06T00:59:41Z
case created โ Released votable artifact with the rare harness-controlled evaluation angle squarely on Scott's harness interests, though traction is currently minimal.
Decision trace
- 10-08 16:42review_screenjev screen: no material development (noul=0.19)
- 10-06 16:21sensor_dirtycomment_update
- 10-06 12:06groundIzolight independently built the public instrument for Scott's model-plus-harness position โ blind A/B voting with the model held constant across the exact harnesses Scott runs (pi, opencode, cod
- 10-06 11:59createReleased votable artifact with the rare harness-controlled evaluation angle squarely on Scott's harness interests, though traction is currently minimal.