HarnessEval is presented as a systematic evaluation of AI code-review configurations, comparing specialist-reviewer harnesses with one-shot prompting using the same models. Its publisher claims the harnesses found 1.6× as many verified bugs in 39 of 42 comparisons, but used roughly 10× more tokens and produced more unsupported findings. The supplied search results discuss similar multi-reviewer workflows, independent verification, and cost per verified outcome, but they do not directly identify HarnessEval, establish who operates it, or independently corroborate its reported figures or Dave Sifry’s role.
The claimed held-model-constant result provides a direct, quantitative test of Scott’s Micro-Agents Architecture and Evaluation-Driven Development: specialist decomposition reportedly raises verified-bug recall, while the 10× token burn and additional unsupported findings pressure his requirements for independent verification and cost per verified outcome. This is a strong dated-receipts and harness-design opportunity, although the figures and publisher identity remain independently uncorroborated in the supplied evidence.
ip:framework.micro-agents-architectureip:concept.evaluation-driven-developmentip:concept.mechanically-different-verifiersip:concept.ai-unit-economicsip:concept.token-disciplineip:source.security-reviewer-method-ebookradar:frontierharness-17x-cost-variationradar:ship-harness-benchradar:cross-model-code-review-validationradar:ai-to-ai-pr-reviewradar:concept.agent-harnessesradar:concept.coding-agent-evaluation
queries asked of Scott's wikis
- specialist agents versus one-shot prompting
- coding-agent harness evaluation and ablations
- AI code review verified recall and false positives
- cost per verified outcome for agent workflows
- parallel reviewers and independent verification
- token budgets versus software quality
now 0 pts/hpeak 27 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 578h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
2026-10-05T19:00:26Z
ReviewBench — an open benchmark for AI code review, reported as GitHub's official one — joins Code Review Bench as a second shared measurement substrate: it does not corroborate the still-publisher-only 1.6× recall claim, but it makes independent evaluation of that claim a routine exercise rather than a waiting game, the first forward-looking shift in the case's meaning since the precision-side remedy landed. Heat stays low despite the magnitude-valve spread reading because the spread is wide but static — the loudest object is the values-essay's spent front-page peak (case-wide 0.17 pts/h, 50th percentile, steady at ~436h) and the newest additions are single-digit, zero-comment posts.
2026-10-05T17:32:20Z
evidence attached: hn.story.49967574 — GitHub's official open benchmark for AI code review supplies the measurement substrate for that exact domain, materially contextualising the specialist-reviewer-harness case.
2026-10-03T07:17:02Z
The precision side of the economics gained its first independent implementation result: a third-party Claude Code skill that verifies each review comment as a claim against the code cut 34% of noise while keeping 93% of real bugs (F1 35.2→40.4) on a human-labelled bench — corroborating the noise premise from the precision direction and turning 'more unsupported findings to triage' from deadweight cost into a tractable, separately-measured problem. The headline 1.6× recall claim remains publisher-only. Heat returns to low: the magnitude-valve spread reading reflects the values-essay's spent front-page peak (now ~0.17 pts/h, 35th percentile, steady at ~377h age), and the newest cross-platform addition is a 1-point post — the periphery of the claim itself is not expanding.
2026-10-03T06:30:14Z
evidence attached: reddit.post.1wweuil — Agentic verification filtering of AI code-review comments (kept 93% of real bugs, cut 34% of noise, F1 35.2→40.4) materially contextualises the AI-code-review reliability case from the complementary precision direction.
2026-09-28T06:39:48Z
Substance unchanged — the benchmark claim remains publisher-only with no independent reproduction; what moved is attention. The attached values-essay ('more to code review than detection') surged 23→112 pts / 56 comments at the 97.5th peer percentile with accelerating momentum, turning the 'verified-bug recall may not measure review's purpose' caveat from a footnote into a live community debate. Heat rises low→medium because the labels under-rate the numbers and the case's periphery is genuinely expanding; it stops short of high because the acceleration is on the adjacent essay, not HarnessEval itself, and the sampled comments never engage the benchmark.
2026-09-27T22:39:49Z
The newly attached front-page essay reframes the debate at the values level — code review exceeds automatable detection — which qualifies HarnessEval's verified-bug-recall success metric even if its numbers hold, but it neither corroborates nor contradicts the benchmark. The case's meaning is essentially unchanged: a publisher-only quantitative claim with anecdotal support for the underlying noise problem, no independent reproduction, and near-zero current engagement (0 pts/h, 18.8th percentile at ~250h age); the essay's 17→23 point drift is engagement, not substance.
2026-09-27T22:24:43Z
evidence attached: hn.story.49857281 — Front-page essay arguing code review exceeds automatable detection materially contextualizes the value and limits of agent code-review harnesses.
2026-09-24T00:55:57Z
Practitioner discussion independently reinforces that open-ended AI review generates noisy findings and that structured passes may help, but it does not validate HarnessEval’s measured gains. The benchmark merits watching while its currently quiet reception and lack of reproduction argue against higher heat or maturity.
2026-09-22T14:23:55Z
evidence attached: reddit.post.1wna77u — The report gives anecdotal field evidence that unconstrained coding-agent review produces many false positives, contextualizing the harness-versus-one-shot review tradeoff.
2026-09-22T02:32:25Z
grounded: converges/high — The claimed held-model-constant result provides a direct, quantitative test of Scott’s Micro-Agents Architecture and Evaluation-Driven Development: specialist d
2026-09-22T02:29:27Z
origin walked (codex/luna, conf 0.98): anchor hn.story.49796016 -> echo.paper.b4448eabad by Dave Sifry
2026-09-22T02:27:55Z
case created — The published comparison provides specific same-model results and an explicit quality-versus-token-cost tradeoff for coding-agent harnesses.