The supplied case attributes Ship Harness Bench to grigio and presents it as a benchmark comparing agent harnesses with the prompt and model held constant. None of the supplied web snippets directly identifies this project or verifies its creator, release, methodology, or results; the similarly named Harness-Bench arXiv result cannot be assumed to be the same benchmark. Other snippets report performance differences when harnesses change around a fixed model, supporting the motivation for such a comparison but not establishing that Ship Harness Bench reliably isolates harness effects.
The proposed comparison repeats Scott’s Model-Plus-Harness Benchmark Unit position and his Trace-backed agent comparison practice; the radar already tracks same-model harness comparisons in FrontierHarness, though no supplied hit establishes that it tracks Ship Harness Bench itself. Without verified methodology, results, or evidence of consequential adoption, this adds no demonstrated basis for changing Scott’s harness choices or evaluation practice.
ip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisonradar:frontierharness-17x-cost-variationradar:concept.agent-harnessesradar:concept.coding-agent-evaluation
queries asked of Scott's wikis
- coding agent harness selection and evaluation
- model versus scaffold performance attribution
- controlled agent benchmarks reproducibility environment resets
- agent context strategy tool design retry loop control
- coding agent success rate token cost latency tradeoffs
2026-09-11T11:32:37Z
No new evidence links back to Ship Harness Bench itself; the case remains an unverified proposal with no supplied methodology or results, while surrounding comment refreshes on adjacent posts (Stellar, runtime-benchmark discussion) only reiterate the general demand for harness evaluation. Nothing in this window advances the specific hypothesis, and no further evidence is expected within horizon.
2026-09-11T08:22:43Z
evidence attached: reddit.post.1wd99iw — This directly reinforces the open hypothesis that agent runtimes need apples-to-apples evaluation separate from model benchmarks.
2026-09-10T21:42:21Z
Stellar adds a concrete lightweight agent-core lead, but its reported Harness-Bench run is not established as using Ship Harness Bench and supplies no controlled performance advantage. The attachment does not advance this case beyond an unverified comparison proposal; adjacent implementations should not substitute for its own methodology or results.
2026-09-10T21:22:58Z
evidence attached: reddit.post.1wcvgc0 — A concrete small agent-core implementation reports a recorded Harness-Bench result, materially informing whether lightweight transparent harnesses can compete.
2026-09-09T17:25:03Z
The Claude–Codex thread adds an opinion favoring validation tools over agent review, not measured evidence about Ship Harness Bench. Adjacent discussion continues to illustrate the evaluation problem without establishing this benchmark’s methodology or decision-useful results; hourly reconsideration is not warranted.
2026-09-08T11:27:14Z
The refreshed discussion adds an anecdotal OpenCode reliability complaint and a linked GUI–harness integration, both belonging to the separate Three.js experiment rather than Ship Harness Bench. Neither establishes this benchmark’s methodology or comparative results; repeated adjacent discussion is not advancing the case.
2026-09-08T08:31:30Z
The refreshed comments add no substantive evidence beyond already-known methodological cautions and subjective preferences in the separate Three.js experiment. Ship Harness Bench remains an unverified controlled-comparison proposal; adjacent discussion does not establish its methodology, results, or usefulness for Scott’s harness decisions.
2026-09-08T07:33:55Z
The refreshed discussion adds a dependency-version concern to the separate Three.js experiment, not a measured harness effect or evidence about Ship Harness Bench. Adjacent evaluation chatter remains repetitive rather than corroborating; this benchmark still lacks supplied methodology and decision-useful results.
2026-09-08T06:31:33Z
New comments on the separate Three.js experiment identify run-to-run variance as an unresolved confound: apparent harness differences need repeated trials before they inform tooling choices. These are methodological questions, not demonstrated failures, and neither corroborate nor disprove Ship Harness Bench.
2026-09-08T05:26:10Z
New comments on the separate Three.js comparison offer subjective preferences for Qwen/OpenCode and highlight visual presentation, result visibility and missing cost information, rather than establish a controlled harness advantage. They do not connect that experiment to Ship Harness Bench, whose methodology and decision-useful results remain unverified.
2026-09-08T04:23:40Z
The Three.js comparison adds another independent experiment in harness evaluation, but its title alone establishes neither a controlled harness effect nor any connection to Ship Harness Bench. This remains evidence of interest in the broader evaluation problem, not validation of this benchmark or a basis for changing Scott’s tooling choices.
2026-09-08T04:22:00Z
evidence attached: hn.story.49605433 — Independent testing of model and harness combinations on a fixed Three.js task materially informs whether harnesses drive coding-agent performance.
2026-09-08T00:26:33Z
The Claude–Codex prototype illustrates demand for harness evaluation, but its author is asking how to benchmark a one-shot demo; it neither validates nor establishes a connection to Ship Harness Bench. This case still lacks supplied methodology or comparative results, and adjacent projects should not count as corroboration.
2026-09-08T00:22:22Z
evidence attached: reddit.post.1wa8sfr — The author is explicitly building a benchmark to compare a multi-agent coding harness, directly bearing on the open harness-evaluation episode.
2026-09-07T11:28:04Z
The new GLM 5.3 comparison supplies a separate lead on same-model harness evaluation, not corroboration of Ship Harness Bench: no connection between the projects is established. Without supplied results or methodology, it supports the broader motivation but does not establish this benchmark's usefulness for Scott's tooling decisions.
2026-09-07T11:22:47Z
evidence attached: hn.story.49596553 — This is a directly relevant independent harness comparison that can materially update the open case about measuring harness effects with a common model.
2026-09-06T10:28:38Z
No substantive delta changes the interpretation: Ship Harness Bench remains a proposed controlled harness comparison, without supplied results or runnable methodology. The GitHub echo repeats the author's premise rather than independently establishing the benchmark or its usefulness.
2026-09-06T10:27:19Z
grounded: known/low — The proposed comparison repeats Scott’s Model-Plus-Harness Benchmark Unit position and his Trace-backed agent comparison practice; the radar already tracks same
2026-09-06T10:25:21Z
case created — A linked benchmark artifact establishes a bounded episode distinct from HarnessOpt-Bench's optimization evaluation, but the supplied evidence contains no comparative results.