2026-10-11 18:02 UTC

grigio presents Ship Harness Bench as a benchmark comparing agent harnesses with the prompt and model held constant, potentially allowing builders to distinguish harness effects from model differences when selecting agent tooling.

state: expiredheat: lowuncertainty: highknownscott: lowagent-harnesses coding-agents evaluationgrigio

What is this?

The supplied case attributes Ship Harness Bench to grigio and presents it as a benchmark comparing agent harnesses with the prompt and model held constant. None of the supplied web snippets directly identifies this project or verifies its creator, release, methodology, or results; the similarly named Harness-Bench arXiv result cannot be assumed to be the same benchmark. Other snippets report performance differences when harnesses change around a fixed model, supporting the motivation for such a comparison but not establishing that Ship Harness Bench reliably isolates harness effects.

Why it matters to Scott

The proposed comparison repeats Scott’s Model-Plus-Harness Benchmark Unit position and his Trace-backed agent comparison practice; the radar already tracks same-model harness comparisons in FrontierHarness, though no supplied hit establishes that it tracks Ship Harness Bench itself. Without verified methodology, results, or evidence of consequential adoption, this adds no demonstrated basis for changing Scott’s harness choices or evaluation practice.
ip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisonradar:frontierharness-17x-cost-variationradar:concept.agent-harnessesradar:concept.coding-agent-evaluation
queries asked of Scott's wikis
  • coding agent harness selection and evaluation
  • model versus scaffold performance attribution
  • controlled agent benchmarks reproducibility environment resets
  • agent context strategy tool design retry loop control
  • coding agent success rate token cost latency tradeoffs

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (7) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShip Harness Bench – Same prompt, same model, different harnessesgrigio13
🟧 echo.github ⭐Ship Harness Bench — “Same prompt, same model, different harnesses.”grigio——
🟧 hn10-task GLM 5.3 harness bench: Claude, OpenCode, pi, zcode, Hermes and 3codecapocasa70
🟠 redditClaude + Codex harness
ClaudeAI
stepanokdev04
🟧 hnI tested 10 model/harness combinations on the same Three.js taskalvins8212574
🟠 redditDo agent frameworks need to be large to be useful?
LocalLLaMA
anandesh-sharma010
🟠 redditAre we missing a benchmark for agent runtimes, not just models?
LocalLLaMA
Balance-1212

Interpretation history

Decision trace