Independent use will determine whether Coarena’s community-submitted pairwise evaluations provide a diverse, current, and useful benchmark for computer-use agents beyond static evaluation suites.
state: expiredheat: lowuncertainty: highconvergesscott: mediumcomputer-use-agents ai-benchmarksCoarena
What is this?
Coarena appears to refer to Computer Agent Arena, an open, crowdsourced platform for evaluating computer-use agents on live Ubuntu and Windows desktops. Users compare two agents side by side on real-world tasks and vote for the better result, generating human-preference data for a continuously updated Elo leaderboard; the full platform stack is published through the xlang-ai GitHub organization and associated with an ICLR 2026 paper. The supplied evidence does not identify individual creators, confirm that “Coarena” and Computer Agent Arena are the same name, or establish meaningful independent adoption or superiority over static benchmarks.
Why it matters to Scott
Coarena independently operationalizes Scott’s Reflexive Agent Design and Progressive Evaluation Ladder patterns: real agents act in live environments, humans compare concrete outcomes, and new usage continually supplies evaluation evidence. It could extend those ideas into a public, cross-model benchmark and provide a useful external test case, but independent adoption, trajectory quality, and benchmark validity remain unestablished.
ip:framework.reflexive-agent-designip:concept.progressive-evaluation-ladderip:concept.human-judgmentip:concept.generate-and-judgeradar:concept.agent-evaluationradar:concept.agent-benchmarksradar:concept.computer-useradar:computer-anthology-terminal-benchmark
queries asked of Scott's wikis
- crowdsourced evaluation for computer-use agents
- pairwise human preference versus static benchmarks
- live-environment agent evaluation
- benchmark contamination and continuously updated evals
- Elo leaderboards for agent systems
- production evaluation of agent trajectories
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-08-09T21:31:44Z
No independent usage, implementation, or benchmark-quality evidence emerged during the launch window, so the episode has faded without validating Coarena’s broader utility. A future adoption or validity result should open a new episode.
2026-08-07T20:30:38Z
Re-evaluation adds no independent adoption, implementation, or benchmark-quality evidence; Coarena remains a promising but creator-asserted evaluation artifact whose broader utility is unproven.
2026-08-07T19:28:30Z
grounded: converges/medium — Coarena independently operationalizes Scott’s Reflexive Agent Design and Progressive Evaluation Ladder patterns: real agents act in live environments, humans co
2026-08-07T19:24:24Z
case created — Coarena is a usable new benchmarking artifact, but participation, task diversity, and comparative value remain unproven.
Decision trace
- 08-10 07:31expireNo independent usage, implementation, or benchmark-quality evidence emerged during the launch window, so the episode has faded without validating Coarena’s broader utility. A future adoption or validi
- 08-10 07:31alert_silentThe only change is elapsed time with unchanged launch engagement and no substantive new evidence; there is nothing Scott needs before a normal briefing.
- 08-10 07:31alert_routeThe only change is elapsed time with unchanged launch engagement and no substantive new evidence; there is nothing Scott needs before a normal briefing.
- 08-08 06:30repriceRe-evaluation adds no independent adoption, implementation, or benchmark-quality evidence; Coarena remains a promising but creator-asserted evaluation artifact whose broader utility is unproven.
- 08-08 06:30alert_silentNothing consequential changed since the launch observation, so this can wait for evidence of independent usage, task diversity, or comparative benchmark validity.
- 08-08 06:30alert_routeNothing consequential changed since the launch observation, so this can wait for evidence of independent usage, task diversity, or comparative benchmark validity.
- 08-08 05:28alert_silentCoarena’s launch is a concrete and relevant evaluation experiment, but the only evidence is the creators’ announcement. There is not yet independent usage, trajectory-quality evidence, or benchmark-va
- 08-08 05:28surface_candidateCoarena’s launch is a concrete and relevant evaluation experiment, but the only evidence is the creators’ announcement. There is not yet independent usage, trajectory-quality evidence, or benchmark-va
- 08-08 05:28alert_routeCoarena’s launch is a concrete and relevant evaluation experiment, but the only evidence is the creators’ announcement. There is not yet independent usage, trajectory-quality evidence, or benchmark-va
- 08-08 05:28groundCoarena independently operationalizes Scott’s Reflexive Agent Design and Progressive Evaluation Ladder patterns: real agents act in live environments, humans compare concrete outcomes, and new usage c
- 08-08 05:24createCoarena is a usable new benchmarking artifact, but participation, task diversity, and comparative value remain unproven.