2026-10-11 17:11 UTC

Terminal-Bench-Science’s maintainers claim their benchmark reproducibly measures AI agents on realistic scientific research workflows, enabling capability comparisons beyond synthetic tasks.

state: expiredheat: lowuncertainty: highconvergesscott: mediumscientific-agents agent-benchmarksTerminal-Bench-Science

What is this?

Terminal-Bench Science is an announced extension of Terminal-Bench, maintained through the Harbor Framework repository and Terminal-Bench team, for evaluating AI agents on complex computational workflows drawn from natural-science research. Its organizers are soliciting scientists to contribute tasks, aiming to create comparable terminal-based evaluations and spur “Claude Code / Codex for Science”-style systems. The supplied snippets establish the benchmark’s goals and contribution process, but provide no completed results or independent evidence that it yet measures scientific-agent capabilities reproducibly.

Why it matters to Scott

Terminal-Bench Science independently converges with Scott’s evaluation-driven and model-plus-harness position, and directly touches his trace-backed agent-comparison work by proposing comparable evaluations on representative workflows. It creates a potential dated-receipts and implementation-comparison opportunity, but remains only an announced benchmark without results or independent reproducibility evidence.
ip:concept.evaluation-driven-developmentip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisonradar:bixbench3-biology-agent-workflowsradar:concept.agent-evaluationradar:concept.scientific-ai
queries asked of Scott's wikis
  • real-world agent evaluation harnesses and reproducibility
  • benchmark-driven agent development and Goodhart effects
  • long-horizon terminal agents and error recovery
  • scientific workflow automation with coding agents
  • agent benchmark integrity contamination and continuous validation
  • portable adapters for cross-agent capability comparisons

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnTerminal-Bench-Science: Evaluating AI agents on scientific research workflowsmatt_d11536
🟧 echo.blog ⭐Announces a benchmark for evaluating AI agents on scientific research workflows.Terminal-Bench-Science——

Interpretation history

Decision trace