Autoprompt's publisher claims its released coding skill raised DeepSeek V4 Flash 0731's Terminal-Bench 2.1 success rate from 67.42% to 82.02% in OpenCode, potentially reducing coding-task failures at the expense of longer runs and higher token costs.
state: seedheat: mediumuncertainty: mediumconvergesscott: mediumcoding-agents agent-harnesses coding-benchmarksSpielewoySorosu
What is this?
Autoprompt is an MIT-licensed coding-agent skill published in Spielewoy's GitHub repository, with explicit invocation because it changes workflow, cost, and runtime. The repository reports that DeepSeek V4 Flash 0731 in OpenCode 1.18.7 solved 73 of 89 Terminal-Bench 2.1 tasks with Autoprompt versus 60 without it (82.02% versus 67.42%), reducing failures from 29 to 16, with an expected trade-off of roughly three times the runtime and twice the tokens. These are publisher-reported results; the supplied snippets do not establish independent replication or whether the improvement persists at matched resource budgets, and the repository explicitly says DeepSeek's separate 82.7% result used a noncomparable setup. The snippets do not establish Sorosu's role or details of the announced v2.
Why it matters to Scott
Autoprompt's released skill and publisher-reported improvement converge with Scott's Model-Plus-Harness Benchmark Unit and Skills and Workflows positions, offering a concrete OpenCode intervention to test using his trace-backed comparison approach—not just another abstract harness claim. The reported roughly 3× runtime and 2× token trade-off makes matched-budget and cost-per-success evaluation necessary before adoption; the radar's DeepSeek V4 Flash harness-efficiency page tracks adjacent comparisons, but the supplied hits do not establish prior coverage of Autoprompt itself.
ip:concept.model-plus-harness-benchmark-unitip:concept.skills-and-workflowsip:concept.ai-unit-economicsdev:concept.trace-backed-agent-comparisondev:technology.opencoderadar:deepseek-v4-flash-harness-efficiencyradar:frontierharness-17x-cost-variationradar:stencil-harness-coding-improvementradar:concept.agent-skillsradar:concept.coding-agent-evaluation
queries asked of Scott's wikis
- Agent harness design versus underlying model capability
- Coding-agent evaluation matched token and runtime budgets
- Cost per successful coding task versus token pricing
- Reusable coding skills and explicit workflow activation
- OpenCode projects and coding-agent skill integration
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 721h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p28 vs 519 stories at the 720h mark (now 721h old) — ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-09-11T15:25:36Z
grounded: converges/medium — Autoprompt's released skill and publisher-reported improvement converge with Scott's Model-Plus-Harness Benchmark Unit and Skills and Workflows positions, offer
2026-09-11T15:22:45Z
case created — A linked release and specific benchmark result establish a bounded harness-efficiency claim, without establishing cross-model gains or equivalent v2 performance.
Decision trace
- 09-20 17:22review_dormantscheduled targets exhausted or 28 quiet days
- 09-20 17:22drop_targetsquiet through full ladder or over cap 8
- 09-12 01:25groundAutoprompt's released skill and publisher-reported improvement converge with Scott's Model-Plus-Harness Benchmark Unit and Skills and Workflows positions, offering a concrete OpenCode interv
- 09-12 01:22createA linked release and specific benchmark result establish a bounded harness-efficiency claim, without establishing cross-model gains or equivalent v2 performance.