2026-10-11 16:37 UTC

HarnessEval’s publisher claims specialist-reviewer harnesses found 1.6 times as many verified bugs as one-shot prompting with the same models in 39 of 42 comparisons, potentially improving AI code review at the cost of roughly tenfold token use and more unsupported findings.

state: watchingheat: lowuncertainty: highconvergesscott: highagent-evaluation agent-harnesses code-review inference-economicsHarnessEvalCompound Engineeringmetareview

What is this?

HarnessEval is presented as a systematic evaluation of AI code-review configurations, comparing specialist-reviewer harnesses with one-shot prompting using the same models. Its publisher claims the harnesses found 1.6× as many verified bugs in 39 of 42 comparisons, but used roughly 10× more tokens and produced more unsupported findings. The supplied search results discuss similar multi-reviewer workflows, independent verification, and cost per verified outcome, but they do not directly identify HarnessEval, establish who operates it, or independently corroborate its reported figures or Dave Sifry’s role.

Why it matters to Scott

The claimed held-model-constant result provides a direct, quantitative test of Scott’s Micro-Agents Architecture and Evaluation-Driven Development: specialist decomposition reportedly raises verified-bug recall, while the 10× token burn and additional unsupported findings pressure his requirements for independent verification and cost per verified outcome. This is a strong dated-receipts and harness-design opportunity, although the figures and publisher identity remain independently uncorroborated in the supplied evidence.
ip:framework.micro-agents-architectureip:concept.evaluation-driven-developmentip:concept.mechanically-different-verifiersip:concept.ai-unit-economicsip:concept.token-disciplineip:source.security-reviewer-method-ebookradar:frontierharness-17x-cost-variationradar:ship-harness-benchradar:cross-model-code-review-validationradar:ai-to-ai-pr-reviewradar:concept.agent-harnessesradar:concept.coding-agent-evaluation
queries asked of Scott's wikis
  • specialist agents versus one-shot prompting
  • coding-agent harness evaluation and ablations
  • AI code review verified recall and false positives
  • cost per verified outcome for agent workflows
  • parallel reviewers and independent verification
  • token budgets versus software quality

Measured heat

now 0 pts/hpeak 27 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 578h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-17 14:00⭐ origin echo-reconstructedOriginal technical report by Dave Sifry. It reports a systematic evaluation of AI code review: “harnesses improve true-gold recall over the
Dave Sifry on paper (echo) · attributed from hn.story.49796016
—
09-22 02:12first on hacker news · published · +108.2hVibes vs. Evidence: What delivers AI code review quality
pjf
—
09-22 13:56first on r/ClaudeAI · published · +119.9hAny tips for reviewing pull requests with Claude Code?
jadedOcelot1
—
09-22 02:12amplified on hacker newshn.story.49796016
pjf
peak 3 · 1 comments · 1% of case engagement
09-22 13:56amplified on r/ClaudeAIreddit.post.1wna77u
jadedOcelot1
peak 5 · 8 comments · 2% of case engagement
09-26 15:06amplified on hacker news 👑hn.story.49857281
utiiiD
peak 165 · 117 comments · 94% of case engagement
10-03 05:11amplified on r/ClaudeAIreddit.post.1wweuil
Content-Berry-2848
peak 3 · 5 comments · 1% of case engagement
10-05 17:11amplified on hacker newshn.story.49967574
azhenley
peak 4 · 0 comments · 1% of case engagement
09-22 02:20our radar first saw it · +108.3hdiscovery anchor: hn.story.49796016—
pace: p79 vs 1032 stories at the 336h mark (now 578h old) — ahead of irregular-agentic-weight-self-modification (1.0x), behind dlab-open-source-week (1.0x)

Evidence (6) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnVibes vs. Evidence: What delivers AI code review quality
Retrieved article excerpt

Open article · Retrieved 2026-09-22T02:22:27.036976+00:00

Finding · harness

1.6×

Same model, add a harness: 1.6× verified bugs found by AI code review

A harness runs the model as a team of specialist reviewers instead of a single prompt. [Compound Engineering](https://github.com/EveryInc/compound-engineering-plugin) and
[metareview](https://github.com/dsifry/metareview) beat the same model’s [one-shot prompt](https://dsifry.github.io/harnesseval/REPORT.html#the-one-shot-baseline-prompt-verbatim) in 39 of 42 comparisons. GLM-5.3 running metareview at
low effort gained 2.1×.

The price: about 10× the tokens, and more unsupported findings to check.

AI code review: same model, add a harness (Compound Engineering or metareview) and it found 1.6× the verified bugs, beating one-shot prompting in 39 of 42 same-model comparisons at about 10× the tokens. Data and method:
pjf31
🟧 echo.paper ⭐Original technical report by Dave Sifry. It reports a systematic evaluation of AI code review: “harnesses improve true-gold recall over the Dave Sifry——
🟠 redditAny tips for reviewing pull requests with Claude Code?
ClaudeAI
jadedOcelot158
🟧 hnThere is more to code review than (automatable) detectionutiiiD165117
🟠 redditI built a Claude Code skill that checks AI code review comments against the code. On CodeRabbit's reviews it removed 34% of the noise and kept 93% of real bugs
ClaudeAI
Content-Berry-284835
🟧 hnReviewBench: An open benchmark for AI code reviewazhenley40

Interpretation history

Decision trace