2026-10-11 17:11 UTC

Independent use will determine whether Computer Anthology’s continuously evolving terminal-task family provides durable agent measurements that resist saturation better than static benchmarks.

state: expiredheat: lowuncertainty: highknownscott: mediumagent-benchmarks terminal-agents coding-agents computer-useVetto AIComputer Anthology

What is this?

Computer Anthology is presented in the case as a Vetto AI benchmark family that continuously evolves terminal tasks for evaluating AI agents. The supplied web snippets do not independently document Computer Anthology itself or establish results from outside adopters; they only support the broader motivation that static, manually curated benchmarks can saturate or diverge from authentic workflows, while terminal and long-horizon evaluations test sustained execution, recovery, and maintainability. Whether Computer Anthology delivers more durable measurements therefore remains an unverified hypothesis in the supplied material.

Why it matters to Scott

Scott already argues for repeatable, contamination-aware evaluation of agents on authentic tool-using paths in “Evaluation-Driven Development,” “Benchmarking the Wrong Unit,” and “Reflexive Agent Design.” Computer Anthology could become a useful external evaluation surface for his Ask terminal agent, but absent independent adoption or results, it currently adds an unverified implementation rather than new evidence about whether continuously evolving benchmarks resist saturation.
ip:concept.benchmarking-the-wrong-unitip:concept.evaluation-driven-developmentip:framework.reflexive-agent-designdev:project.askradar:concept.agent-benchmarksradar:concept.agent-evaluationradar:concept.benchmark-integrityradar:concept.agent-harnesses
queries asked of Scott's wikis
  • evolving benchmarks and benchmark saturation
  • coding-agent evaluation on authentic workflows
  • terminal-agent harnesses and task verification
  • long-horizon agent reliability and recovery
  • benchmark contamination and durable measurements
  • agent evaluation beyond pass/fail tests

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnComputer Anthology: A continuously evolving benchmark family for AI agentsrigelbm2911
🟧 echo.blog ⭐Introduces Computer Anthology as a continuously evolving benchmark family for AI agents performing terminal tasks.Vetto AI——

Interpretation history

Decision trace