2026-10-11 17:10 UTC

Independent replication will determine whether Stencil's harness-only changes reproducibly improve coding performance across 15 different LLMs as claimed.

state: expiredheat: lowuncertainty: highconvergesscott: highagent-harnesses coding-agents evaluation

What is this?

Stencil’s reported result concerns “hashline,” a code-editing format that tags lines with content hashes so models can specify edits without reproducing exact source text. The supplied article snippet claims this harness-only change improved coding results across 15 LLMs—by as much as 61.6 percentage points—and reduced token usage by 20–30%; the case attributes the earliest primary artifact to Can Bölük’s oh-my-pi commit adding a hashline edit mode. The snippets offer broader support for harness effects, but they do not establish an independent replication of Stencil’s specific 15-model experiment, and they do not clearly identify who or what Stencil is.

Why it matters to Scott

Stencil's claim that a harness-only change (hashline edit format) reproducibly lifts coding performance across 15 different LLMs is an independent, empirical instance of Scott's own thesis that agentic capability is a property of model-plus-harness, not weights alone (Model-Plus-Harness Benchmark Unit) — and it bears directly on his Benchmarking the Wrong Unit critique and his own trace-backed agent-comparison practice. This is a dated-receipts opportunity: a third party arriving at, and quantifying, a position Scott has already staked out and built tooling around.
ip:concept.model-plus-harness-benchmark-unitip:concept.benchmarking-the-wrong-unitdev:concept.trace-backed-agent-comparisonip:concept.surgery-problemip:concept.deterministic-verification-before-assertionradar:benzi-repository-map-harness
queries asked of Scott's wikis
  • coding-agent edit formats and patch reliability
  • agent harness versus base-model capability
  • cross-model harness evaluation methodology
  • content-addressed code editing and line hashes
  • coding-agent benchmark harness confounds
  • token-efficient tool interfaces for code agents

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnWe improved 15 LLMs at coding in one afternoon. Only the harness changedlatchkey10
🟧 echo.github ⭐The earliest primary artifact is Can Bölük’s oh-my-pi commit introducing the hashline edit mode: “added hashline edit mode with xxHash64 intCan Bölük——
🟧 hnShow HN: Auto-train the harness, not the LLM. cross-model, cross-benchmark gainsmegadragon930
🟠 redditOne Claude Code skill pushed DeepSeek V4 Flash from 67.42% to 82.02%
ClaudeAI
Sorosu18940

Interpretation history

Decision trace