2026-10-11 18:01 UTC

scientific-agents

band: coolmomentum: stable score: 0.051
temperature history

Episodes (2)

Independent replication will determine whether Anthropic’s Claude-assisted protein-design workflow materially improves wet-lab success rates over conventional human-led design.
expiredknownscott: medium
Terminal-Bench-Science’s maintainers claim their benchmark reproducibly measures AI agents on realistic scientific research workflows, enabling capability comparisons beyond synthetic tasks.
expiredconvergesscott: medium

Trajectory notes