2026-10-11 16:38 UTC

Gamow Labs claims its released LabBench — 20 held-out wet-lab decision tasks built from real drug-discovery and genomics records — shows five frontier agents pass only ~40% of decision criteria, fail every criterion on which experiment should come first, and recover on failed decisions only when a one-sentence attention redirect is appended, and adoption by AI-for-science evaluators would establish it as the reference benchmark for whether agents can decide the next wet-lab experiment.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumagent-evaluation ai-for-science benchmarksGamow LabsBen ReillyLeah DamonDaniel McKinnon

What is this?

Per the case, Gamow Labs (with Ben Reilly, Leah Damon, and Daniel McKinnon named) has released LabBench, 20 held-out wet-lab decision tasks built from real drug-discovery and genomics records, claiming five frontier agents pass only ~40% of decision criteria (182 of 406 for the best), fail every experiment-ordering criterion, and recover on failed decisions only when a one-sentence attention redirect is appended. The supplied web results never mention LabBench or Gamow Labs, so the release and its numbers rest on first-party claims only; they do show a crowded adjacent band — Scale's DrugDiscoveryBench (82 tasks, top agent 51.6), Insilico's DDD Benchmark-as-a-Service, Elicit's BioDecisionBench, AIRS-Bench — all targeting agentic scientific decision-making. Notably, Scale's expert-playbook recovery experiment lands on the same shape of conclusion as LabBench: agents solve most tasks when given external structure, locating the primary gap in unguided high-level planning rather than tool execution.

Why it matters to Scott

Converges with his attention-interrupt canon: a third-party benchmark independently finds frontier agents fail at choosing the next wet-lab experiment yet recover on a one-sentence attention redirect — dated evidence that recovery lives in context steering rather than model capability, exactly the Prompt-Interrupt Architecture / Attention-Routing / Attention-Budget position, and the same-agents-pass-only-when-the-harness-changes result is external support for his Model-Plus-Harness Benchmark Unit. It also quantifies the experiment-selection gap his autonomous-science radar cases (Anthropic's robot lab, Roche) presuppose is closing; but first-party-only claims, minimal traction, and a wet-lab domain he doesn't operate in keep it at watch-level unless independent adoption confirms the reference-benchmark role.
ip:framework.prompt-interrupt-architectureip:concept.attention-routingip:concept.attention-budgetip:concept.model-plus-harness-benchmark-unitip:framework.cognition-dimension-ladderradar:concept.ai-for-scienceradar:concept.scientific-agentsradar:concept.agent-evaluationradar:concept.ai-benchmarksradar:concept.benchmark-integrityradar:concept.context-engineeringradar:terminal-bench-science-workflowsradar:anthropic-preclinical-robot-lab
queries asked of Scott's wikis
  • agent harness evaluation loop design
  • coding agents planning versus execution gap
  • attention redirect prompt steering agent recovery
  • autonomous science experiment selection agent
  • first-party benchmark as product moat
  • agent memory experiment decision records

Measured heat

now 0 pts/hpeak 2 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 338h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-27 14:00⭐ origin echo-reconstructed'LabBench is 20 held-out tasks built from real wet-lab records'; 'The best agents pass two in five criteria', 182 of 406 criteria passed by
Gamow Labs — Ben Reilly, Leah Damon, Daniel McKinnon on blog (echo) · attributed from hn.story.49903060
—
09-30 00:59first on hacker news · published · +59.0hLabBench: Can AI agents decide what experiment to run next?
wardbradt
—
09-30 00:59amplified on hacker news 👑hn.story.49903060
wardbradt
peak 4 · 0 comments · 98% of case engagement
09-30 03:21our radar first saw it · +61.4hdiscovery anchor: hn.story.49903060—
pace: p36 vs 1032 stories at the 336h mark (now 338h old) — ahead of agentgate-signed-agent-receipts (1.3x), behind agent-memory-add-search-evaluation (0.8x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnLabBench: Can AI agents decide what experiment to run next?
Retrieved article excerpt

Open article · Retrieved 2026-09-30T03:31:34.958727+00:00

September 28, 2026 · Ben Reilly, Leah Damon, Daniel McKinnon

# LabBench: Can AI agents decide what experiment to run next?

From our preprint, [LabBench: Benchmarking AI wet lab experimental design and decision-making for AI Agents](https://gamowlabs.com/assets/labbench-preprint.pdf) (PDF). [Partner with us](https://gamowlabs.com/labbench-benchmarking-ai-wet-lab-decisions.html#partner).

Accelerating biological experimentation relies on quickly deciding which experiment comes next. Picking the test that can fastest invalidate or support a hypothesis is a skill only human experts perform reliably today. To scale discovery past human capability, we have to evaluate and improve AI agents on it.

**LabBench** is 20 held-out tasks built from real wet-lab records in drug discovery and genomics. Each is a snapshot in time: the agent gets the records that existed at a decision point, and the lab’s interpretation and decision are withheld. It must commit to the next step, graded against 20–22 binary criteria tied to what the lab actually decided.

From lab archive to graded task

1. **Mine**Find recorded decisions and their evidence.
2. **Assemble**Give the records; withhold the decision.
3. **Rubric**20–22 binary criteria from the true answer.
4. **Harden**Rework tasks a solver passes trivially.
5. **Review**Biologists verify every task.

The tasks come from real lab data unlikely to be recalled from the open internet. We expect the agent to navigate the evidence as a real researcher would, without the brief pointing to what matters.

## Results

Five frontier agents each ran once per task in their vendor’s harness.

The best agents pass two in five criteria

Mean share of criteria passed; line is the 95% interval.

GPT-6 Astra40.8%34–48

Claude Opus 5.540.5%33–49

Grok 4.729.6%23–37

Muse Spark 1.328.2%23–34

Gemini 3.8 Flash18.4%13–24

0255075100%

GPT-6 Astra and Claude Opus 5.5 tie, but Astra took a median 5 minutes per task to Opus’s 35.

Same score, seven times faster

Score against median minutes per task (log scale).

0%20%40%60%

GPT-6 Astra40.8% · 5 minClaude Opus 5.540.5% · 34.6 minGrok 4.729.6% · 14.2 minMuse Spark 1.328.2% · 4.5 minGemini 3.8 Flash18.4% · 13.3 min

35102040

GPT-6 Astra 40.8%, 5 minClaude Opus 5.5 40.5%, 35 minGrok 4.7 29.6%, 14 minMuse Spark 1.3 28.2%, 4.5 minGemini 3.8 Flash 18.4%, 13 min

### Agents fail the same criteria

182 of 406 criteria were passed by no agent; 31 by all five. What one frontier agent misses, the others usually miss too.

45% of criteria were passed by no agent

Criteria by number of agents passing them.

**182**

**55**

**40**

**45**

**53**

**31**

none1234all 5

### They interpret; they don’t choose

Agents excel at interpreting previous experiments: saying what a measurement is, declining an overclaim, reconstructing an analysis. Criteria that require choosing, committing or ranking pass 21% of the time, against 47% for identifying what something is. No agent passed any of the 13 criteria on which experiment should come first.

Choosing is the hardest move

Pass rate by what the criterion asks; “none” = share no agent passed.

Identify what something is47%none: 32%

Reconstruct or explain38%none: 37%

Qualify a claim33%none: 39%

Integrate evidence31%none: 38%

Quantify or read out30%none: 46%

**Choose, commit or rank**21%none: 64%

0255075100%

### Where agents differ

Astra leads on core decisions (52%), Opus on evidence integration (53%). On experiment design, the best agent passes 9%.

No agent can design the next experiment

% of each skill’s criteria passed (criteria count beside skill).

GPT-6  
AstraClaude  
Opus 5.5Grok  
4.7Muse  
Spark 1.3Gemini  
3.8 FlashMethod reconstruction 425043383338Measurement semantics 435644403328Evidential status 415149443912Comparison & baseline 404850323530Confounds & data quality 273752443315Scope & generalization 294852243121Quantitative reading 554542403116Core decision 235243302617Evidence integration 364753223611Experiment design 7039033



Deciding is harder than analyzing

% of criteria passed, by theme and task type.

GPT-6 AstraClaude Opus 5.5Grok 4.7Muse Spark 1.3Gemini 3.8 Flash

Nascent transcriptionn=42

Gene regulationn=60

3D genomen=62

Drug responsen=101

Hit-to-lead chemistryn=40

Target validationn=41

Assay developmentn=60

Analyze tasksn=261

Review tasksn=84

Design tasksn=61

0204060%

## The knowledge is latent

A common issue with hard benchmarks is unreasonable criteria that make tasks effectively impossible. We tested this directly. On five core decisions no agent passed, we appended one sentence pointing at evidence GPT-6 Astra already had, without stating the answer. It passed all five.

One sentence flips the decision

GPT-6 Astra’s task score. The core decision flipped from fail to pass in all five.

Benchmark runWith one added sentence

Next two experimentsdrug response20% → 75%

Chromosome model review3D genome23% → 64%

Candidate coding transcriptsnascent transcription50% → 82%

Motif evidence across datasetsgene regulation35% → 60%

Library preparation choiceassay development35% → 40%

0255075100%

One run per hinted task.

#### Drug response

[Gene-set enrichment running score for TNF-alpha signaling via NF-kappa-B at 15 minutes under four death-inducing treatments; one treatment rises far above the other three.](https://gamowlabs.com/assets/labbench-nfkb-signature.png)

NF-κB signaling at 15 minutes under four death inducers, from the task’s records (compound names withheld).

Evidence  
A compound triggers the strongest early NF-κB signature of four death inducers; nothing tests whether it drives death.

Agents  
Four of five designed the NF-κB test, then ranked it second.

Hint  
“Rank first the experiment that could tell a cause from a bystander.” Score: 4/20 → 15/20.

#### Assay development

[Fragment-size trace peaking at 210 bp with a smoothly declining tail.](https://gamowlabs.com/assets/labbench-trace-5ul-2min.png)

5 µL Tn5, 2 min: peak 210 bp, clean tail. **The lab’s choice.**

[Fragment-size trace peaking at 238 bp with a broad shoulder of large fragments extending past 1000 bp.](https://gamowlabs.com/assets/labbench-trace-2p5ul-5min.png)

2.5 µL Tn5, 5 min: peak 238 bp, broad large-fragment shoulder. Every agent’s choice.

Agents  
All five chose the larger 238 bp peak and ignored the shoulder of uncut DNA.

Hint  
“Judge each tagmentation trace by its whole size distribution…” The agent chose the lab’s condition.

The hints add no biology; they redirect attention. The models have the knowledge but do not reliably recall it.

## A reasoning bottleneck

Frontier agents know enough biology to work alongside expert biologists when experimental design stays with humans. Their ability to design experiments independently is weak: they fail to commit to an experiment and to discriminate the best one from the alternatives.

In real biological experimentation, wall-clock time is irreducible. Cells grow, differentiate and respond on their own schedules, and each experiment consumes weeks and material. Autonomous and correct experimental design is arguably the single most important capability for making AI agents superhuman at these tasks.

LabBench measures this directly. Improved results should indicate agents that can carry out supervised, autonomous wet-lab experimentation, and eventually run full programs themselves. To get there, models must develop a more cohesive world model of what experiments cost and of the specific evidence that would count against a hypothesis.

## Work with us

Gamow Labs is uniquely equipped to produce data at the frontier of AI and biological experimentation. We are capable of producing the aforementioned tasks at scale. Reach out.

[email protected]

Full methods are in the [preprint](https://gamowlabs.com/assets/labbench-preprint.pdf). We’re also hiring.

Partner with us
[Preprint (PDF)](https://gamowlabs.com/assets/labbench-preprint.pdf)
[Open roles](https://gamowlabs.com/careers.html)
wardbradt40
🟧 echo.blog ⭐'LabBench is 20 held-out tasks built from real wet-lab records'; 'The best agents pass two in five criteria', 182 of 406 criteria passed by Gamow Labs — Ben Reilly, Leah Damon, Daniel McKinnon——

Interpretation history

Decision trace