2026-10-11 16:38 UTC

Autoprompt's publisher claims its released coding skill raised DeepSeek V4 Flash 0731's Terminal-Bench 2.1 success rate from 67.42% to 82.02% in OpenCode, potentially reducing coding-task failures at the expense of longer runs and higher token costs.

state: seedheat: mediumuncertainty: mediumconvergesscott: mediumcoding-agents agent-harnesses coding-benchmarksSpielewoySorosu

What is this?

Autoprompt is an MIT-licensed coding-agent skill published in Spielewoy's GitHub repository, with explicit invocation because it changes workflow, cost, and runtime. The repository reports that DeepSeek V4 Flash 0731 in OpenCode 1.18.7 solved 73 of 89 Terminal-Bench 2.1 tasks with Autoprompt versus 60 without it (82.02% versus 67.42%), reducing failures from 29 to 16, with an expected trade-off of roughly three times the runtime and twice the tokens. These are publisher-reported results; the supplied snippets do not establish independent replication or whether the improvement persists at matched resource budgets, and the repository explicitly says DeepSeek's separate 82.7% result used a noncomparable setup. The snippets do not establish Sorosu's role or details of the announced v2.

Why it matters to Scott

Autoprompt's released skill and publisher-reported improvement converge with Scott's Model-Plus-Harness Benchmark Unit and Skills and Workflows positions, offering a concrete OpenCode intervention to test using his trace-backed comparison approach—not just another abstract harness claim. The reported roughly 3× runtime and 2× token trade-off makes matched-budget and cost-per-success evaluation necessary before adoption; the radar's DeepSeek V4 Flash harness-efficiency page tracks adjacent comparisons, but the supplied hits do not establish prior coverage of Autoprompt itself.
ip:concept.model-plus-harness-benchmark-unitip:concept.skills-and-workflowsip:concept.ai-unit-economicsdev:concept.trace-backed-agent-comparisondev:technology.opencoderadar:deepseek-v4-flash-harness-efficiencyradar:frontierharness-17x-cost-variationradar:stencil-harness-coding-improvementradar:concept.agent-skillsradar:concept.coding-agent-evaluation
queries asked of Scott's wikis
  • Agent harness design versus underlying model capability
  • Coding-agent evaluation matched token and runtime budgets
  • Cost per successful coding task versus token pricing
  • Reusable coding skills and explicit workflow activation
  • OpenCode projects and coding-agent skill integration

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 721h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-11 15:22 (minted)⭐ origin echo-reconstructedThe publisher links this repository for Autoprompt, reports a v1.0 Terminal-Bench 2.1 improvement from 67.42% to 82.02%, and announces v2 su
Spielewoy on github (echo) · attributed from reddit.post.1wdi7vr · published time unknown
—
09-11 14:41first on r/ClaudeAI · published · lag ?A Claude Code skill pushed DeepSeek V4 Flash from 67.42% to 82.02%
Sorosu
—
09-11 14:41amplified on r/ClaudeAI 👑reddit.post.1wdi7vr
Sorosu
peak 1 · 1 comments · 96% of case engagement
09-11 15:20our radar first saw it · lag ?discovery anchor: reddit.post.1wdi7vr—
pace: p28 vs 519 stories at the 720h mark (now 721h old) — ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditA Claude Code skill pushed DeepSeek V4 Flash from 67.42% to 82.02%
ClaudeAI
Sorosu01
🟧 echo.github ⭐The publisher links this repository for Autoprompt, reports a v1.0 Terminal-Bench 2.1 improvement from 67.42% to 82.02%, and announces v2 suSpielewoy——

Interpretation history

Decision trace