Failure Map's creator claims the released archive of 20,168 Python boundary-case repair tasks across 254 categories โ with explicit contracts and executable boundary checks โ becomes an adopted evaluation resource for local coding models' debugging; sustained external use confirms it, fading into an unadopted personal release refutes it.
state: seedheat: lowuncertainty: mediumconvergesscott: mediumagent-evaluation local-coding-models benchmark-datasetsfailuremap-f
What is this?
Failure Map is a first-party release (creator: failuremap-f) of an archive of 20,168 Python repair tasks across 254 categories, each pairing a boundary-case bug with explicit contracts and executable boundary checks, positioned as an evaluation resource for local coding models' debugging. The supplied web snippets contain no third-party coverage of Failure Map itself โ the release's existence and framing rest on the creator's own post โ so zero external traction currently remains unverified either way. The snippets do confirm a live adjacent territory: a 2026 empirical study of locally deployed LLMs on bug repair reports only 43โ45% accuracy with many partially-correct responses, and multiple benchmark-construction efforts for completion, repair, and fault localization are active. The hypothesis is thus an adoption bet in an active but crowded evaluation space, falsifiable by tracking whether anyone outside the creator runs the corpus.
Why it matters to Scott
The corpus's explicit-contracts-plus-executable-boundary-checks design independently instantiates the acceptance doctrine Scott's canon already argues โ challenger-never-arbiter's 'no model judge, replay against shared measures', validation-gated acceptance, and review-until-clear's pre-written answer key โ and it targets exactly the local-model debugging territory he runs on gamepc/Ollama, making it an eval asset he could actually pull down and use against his own agents rather than a mere illustration. It stays medium rather than high: a zero-traction, unknown-creator release in a crowded repair-corpus field (Goldset already runs the sibling adoption bet) is an adoption watch, not a challenge to anything load-bearing or a dated-receipts moment against a consequential party.
ip:framework.challenger-never-arbiterdev:concept.review-until-clear-loopdev:concept.validation-gated-llm-extractiondev:project.gamepcradar:goldset-python-repair-corpusradar:concept.local-inferenceradar:concept.benchmarksradar:concept.agent-debugging
queries asked of Scott's wikis
- coding agent eval harness executable checks
- local open-weight model coding evaluation
- benchmark dataset release adoption dynamics
- contract-based verifiable tasks agent testing
- debug repair workflow coding agent tooling
- indie long-tail release traction tracking
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 235h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p8 vs 1188 stories at the 168h mark (now 235h old) โ behind addom-local-coding-harness (0.5x)
Evidence (2) โ โญ canonical anchor
Interpretation history
2026-10-02T00:10:09Z
origin walked (opencode/cheap-glm, conf 0.85): anchor reddit.post.1wvbk9q -> echo.other.8ed28a2858 by Failure Map (self-published project; Reddit poster u/failuremap-f is its creator)
2026-10-01T23:49:30Z
grounded: converges/medium โ The corpus's explicit-contracts-plus-executable-boundary-checks design independently instantiates the acceptance doctrine Scott's canon already argues โ challen
2026-10-01T23:42:00Z
case created โ A concrete first-party release of a large executable repair-task corpus aimed at local-model debugging is a trackable adoption episode despite zero current traction.
Decision trace
- 10-02 10:10promote_anchororigin walk conf 0.85
- 10-02 09:49groundThe corpus's explicit-contracts-plus-executable-boundary-checks design independently instantiates the acceptance doctrine Scott's canon already argues โ challenger-never-arbiter's '
- 10-02 09:42createA concrete first-party release of a large executable repair-task corpus aimed at local-model debugging is a trackable adoption episode despite zero current traction.