Stencil’s reported result concerns “hashline,” a code-editing format that tags lines with content hashes so models can specify edits without reproducing exact source text. The supplied article snippet claims this harness-only change improved coding results across 15 LLMs—by as much as 61.6 percentage points—and reduced token usage by 20–30%; the case attributes the earliest primary artifact to Can Bölük’s oh-my-pi commit adding a hashline edit mode. The snippets offer broader support for harness effects, but they do not establish an independent replication of Stencil’s specific 15-model experiment, and they do not clearly identify who or what Stencil is.
Stencil's claim that a harness-only change (hashline edit format) reproducibly lifts coding performance across 15 different LLMs is an independent, empirical instance of Scott's own thesis that agentic capability is a property of model-plus-harness, not weights alone (Model-Plus-Harness Benchmark Unit) — and it bears directly on his Benchmarking the Wrong Unit critique and his own trace-backed agent-comparison practice. This is a dated-receipts opportunity: a third party arriving at, and quantifying, a position Scott has already staked out and built tooling around.
ip:concept.model-plus-harness-benchmark-unitip:concept.benchmarking-the-wrong-unitdev:concept.trace-backed-agent-comparisonip:concept.surgery-problemip:concept.deterministic-verification-before-assertionradar:benzi-repository-map-harness
queries asked of Scott's wikis
- coding-agent edit formats and patch reliability
- agent harness versus base-model capability
- cross-model harness evaluation methodology
- content-addressed code editing and line hashes
- coding-agent benchmark harness confounds
- token-efficient tool interfaces for code agents
2026-08-22T19:37:57Z
The practical replication window has passed without an inspectable independent test of Stencil’s hashline intervention. Adjacent harness results support the broader model-plus-harness thesis but do not validate this specific 15-model claim, so this episode has faded rather than matured.
2026-08-20T18:36:16Z
The refreshed discussion adds no controlled comparison or independent replication of Stencil’s hashline intervention; it remains repetitive amplification of the broader harness thesis. The specific 15-model claim stays plausible but uncorroborated and cold.
2026-08-20T02:29:30Z
The refreshed comments are repetitive amplification and requests for comparisons, not a controlled result or replication of Stencil’s hashline intervention. The broader harness thesis remains plausible, but this specific claim stays cold and uncorroborated.
2026-08-19T22:39:42Z
The refreshed discussion mainly identifies the existing compute-budget confound: the reported workflow gain used roughly twice the tokens and three times the runtime without a matched baseline. No independent replication of Stencil’s hashline intervention or new controlled result has emerged, so the case remains open but cold.
2026-08-19T18:33:31Z
The DeepSeek workflow result adds another independent example of harness changes shifting coding benchmarks, strengthening the broader model-plus-harness pattern. It does not replicate Stencil’s hashline intervention or validate its 15-model results, so the specific claim remains uncorroborated and cold.
2026-08-19T18:23:40Z
evidence attached: reddit.post.1vst0hm — This is an independent coding benchmark result suggesting a workflow or skill-only harness change can substantially improve DeepSeek V4 Flash performance, directly bearing on harness-only gains.
2026-08-18T01:32:02Z
Another staleness check yields no inspectable replication of the hashline result; the broader harness thesis remains relevant, but this specific claim is still a cold, first-party result.
2026-08-16T01:23:03Z
The staleness check adds no replication, methods, or benchmark results; the second harness claim remains adjacent rather than corroborating Stencil’s hashline result. The case stays open but cold pending an inspectable independent test.
2026-08-14T00:36:44Z
A separate project now claims cross-model, cross-benchmark gains from optimizing the harness, making this a broader pattern worth watching rather than an isolated Stencil claim. Its available evidence does not show methods, results, or replication of Stencil’s hashline intervention, so corroboration remains unearned.
2026-08-14T00:22:32Z
evidence attached: hn.story.49293267 — This is an independent first-party artifact directly testing the open hypothesis that harness-only changes improve coding performance across models.
2026-08-12T07:30:59Z
No new engagement beyond a trivial score tick (1→2, still zero comments); no independent replication or discussion has emerged since grounding. Case remains a single first-party claim plus its own primary commit history.
2026-08-12T07:29:25Z
grounded: converges/high — Stencil's claim that a harness-only change (hashline edit format) reproducibly lifts coding performance across 15 different LLMs is an independent, empirical in
2026-08-12T07:26:21Z
origin walked (codex/luna, conf 0.96): anchor hn.story.49268543 -> echo.github.dba810c09a by Can Bölük
2026-08-12T07:24:14Z
case created — New first-party blog claim of cross-model harness gains, no prior case covers this specific report and it has zero engagement yet.