Terminal-Bench Science is an announced extension of Terminal-Bench, maintained through the Harbor Framework repository and Terminal-Bench team, for evaluating AI agents on complex computational workflows drawn from natural-science research. Its organizers are soliciting scientists to contribute tasks, aiming to create comparable terminal-based evaluations and spur “Claude Code / Codex for Science”-style systems. The supplied snippets establish the benchmark’s goals and contribution process, but provide no completed results or independent evidence that it yet measures scientific-agent capabilities reproducibly.
Terminal-Bench Science independently converges with Scott’s evaluation-driven and model-plus-harness position, and directly touches his trace-backed agent-comparison work by proposing comparable evaluations on representative workflows. It creates a potential dated-receipts and implementation-comparison opportunity, but remains only an announced benchmark without results or independent reproducibility evidence.
ip:concept.evaluation-driven-developmentip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisonradar:bixbench3-biology-agent-workflowsradar:concept.agent-evaluationradar:concept.scientific-ai
queries asked of Scott's wikis
- real-world agent evaluation harnesses and reproducibility
- benchmark-driven agent development and Goodhart effects
- long-horizon terminal agents and error recovery
- scientific workflow automation with coding agents
- agent benchmark integrity contamination and continuous validation
- portable adapters for cross-agent capability comparisons
2026-08-31T15:37:09Z
After repeated checks and 48 hours without artifacts, results, methodology, or independent use, the announcement has faded without substantiating its reproducibility claim. Reopen only if the maintainers release an executable benchmark or external teams report results.
2026-08-29T15:33:15Z
The latest comments repeat anecdotal model comparisons and correctness concerns without adding tasks, methodology, results, or independent reproduction. The benchmark remains a relevant proposal, not yet evidence of reproducible scientific-agent measurement.
2026-08-28T14:41:50Z
The refreshed comments add concerns about correctness and anecdotal model comparisons, but no methodology, released tasks, results, or independent reproduction. Repeated discussion does not change the case from an announced benchmark awaiting substantive artifacts or external use.
2026-08-28T09:31:23Z
The refreshed discussion remains anecdotal model preference and broad enthusiasm, adding no methodology, artifacts, benchmark results, or independent reproduction. The case still represents a relevant benchmark proposal rather than evidence that realistic scientific-agent workflows are being measured reproducibly.
2026-08-28T01:32:46Z
Refreshed comments merely endorse realistic workflow evals or make unsupported model comparisons; they add no artifacts, methodology, results, or independent reproducibility evidence. The case remains an announced benchmark awaiting implementation or external use.
2026-08-28T00:28:02Z
No substantive evidence has appeared beyond the original benchmark announcement; modest engagement does not validate its task design, results, or reproducibility. The case remains a relevant proposal awaiting artifacts or independent use.
2026-08-28T00:26:32Z
grounded: converges/medium — Terminal-Bench Science independently converges with Scott’s evaluation-driven and model-plus-harness position, and directly touches his trace-backed agent-compa
2026-08-28T00:24:34Z
case created — This is a concrete first-party benchmark release addressing a consequential and distinct scientific-agent evaluation gap.