Independent use will determine whether Goldset’s 896 test-verified Python repairs provide a reproducible and decision-useful benchmark or training corpus for code-repair agents.
state: expiredheat: lowuncertainty: highknownscott: lowcoding-agents code-repair evaluationAndy Salvo
What is this?
Goldset is presented as a released corpus of 896 Python bug fixes whose validity was checked by running the associated test suites; the supplied material names Andy Salvo but does not establish his precise role. The surrounding snippets show why independent evaluation matters: public code-repair sets may poorly match real repositories, may be contaminated by model training data, and can reward test-passing patches that maintainers would not merge. The supplied results do not independently document Goldset’s construction, repository coverage, licensing, reproducibility, or performance as either a benchmark or training corpus, so those claims remain open.
Why it matters to Scott
Scott already holds the relevant position: coding-agent corpora require contamination-resistant, reproducible, harness-aware evaluation, with training and held-out evaluation data kept distinct. Goldset is currently only another candidate corpus in a territory already tracked by the radar; without construction details, independent reruns, or evidence that it changes agent comparisons, it does not yet extend or challenge that position.
ip:framework.hidden-gates-frameworkip:concept.future-leakage-ruledev:concept.trace-backed-agent-comparisonradar:concept.coding-agent-benchmarksradar:concept.benchmark-integrityradar:gitskills-agent-skill-dataset
queries asked of Scott's wikis
- coding-agent evaluation on real repositories
- test-passing patches versus merge-worthy repairs
- code-repair benchmark contamination and leakage
- golden datasets for coding-agent evals
- reproducible harnesses for repository-level repair
- benchmark corpora as training data versus evaluation data
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-21T15:36:08Z
After 48 hours, no independent use, reproducibility report, methodological detail, or discussion has emerged; the release has faded without advancing beyond an unvalidated candidate corpus.
2026-08-19T14:37:18Z
No independent adoption, reruns, or methodological detail has appeared; the case remains an unvalidated corpus release rather than evidence of a decision-useful benchmark or training set.
2026-08-19T14:33:18Z
grounded: known/low — Scott already holds the relevant position: coding-agent corpora require contamination-resistant, reproducible, harness-aware evaluation, with training and held-
2026-08-19T14:30:54Z
case created — The repository is a concrete evaluation artifact, but it has not yet attracted independent use or validation.
Decision trace
- 08-22 01:36expireAfter 48 hours, no independent use, reproducibility report, methodological detail, or discussion has emerged; the release has faded without advancing beyond an unvalidated candidate corpus.
- 08-22 01:36alert_silentThe reobservation is unchanged and adds no consequential delta; revive only if an independent rerun, implementation, or benchmark result appears.
- 08-22 01:36alert_routeThe reobservation is unchanged and adds no consequential delta; revive only if an independent rerun, implementation, or benchmark result appears.
- 08-20 00:37repriceNo independent adoption, reruns, or methodological detail has appeared; the case remains an unvalidated corpus release rather than evidence of a decision-useful benchmark or training set.
- 08-20 00:37alert_silentThe reobservation is unchanged and adds no consequential delta; wait for an independent implementation, reproducibility report, or evidence that Goldset changes coding-agent comparisons.
- 08-20 00:37alert_routeThe reobservation is unchanged and adds no consequential delta; wait for an independent implementation, reproducibility report, or evidence that Goldset changes coding-agent comparisons.
- 08-20 00:33alert_silentThe corpus release is established, but the visible evidence adds only its size and test-suite verification. Without construction methodology, contamination controls, repository/harness details, held-o
- 08-20 00:33surface_candidateThe corpus release is established, but the visible evidence adds only its size and test-suite verification. Without construction methodology, contamination controls, repository/harness details, held-o
- 08-20 00:33alert_routeThe corpus release is established, but the visible evidence adds only its size and test-suite verification. Without construction methodology, contamination controls, repository/harness details, held-o
- 08-20 00:33groundScott already holds the relevant position: coding-agent corpora require contamination-resistant, reproducible, harness-aware evaluation, with training and held-out evaluation data kept distinct. Golds
- 08-20 00:30createThe repository is a concrete evaluation artifact, but it has not yet attracted independent use or validation.