DeepSWE-mini creator asankhs claims the released 16-instance subset preserves the full DeepSWE leaderboard's relative model rankings, potentially reducing the cost of routine local coding-agent evaluation without reproducing absolute scores.
state: seedheat: mediumuncertainty: mediumconvergesscott: mediumagent-evaluation coding-agents benchmarkingasankhs
What is this?
DeepSWE is Datacurve’s benchmark for long-horizon coding agents, covering 113 original tasks across 91 repositories and five languages, with isolated environments, executable verifiers, and a fixed mini-swe-agent harness. The case attributes DeepSWE-mini to asankhs, who claims a released 16-task subset preserves the full leaderboard’s relative rankings rather than its absolute scores. The supplied web snippets establish the parent benchmark but do not directly document the mini release, its creator, or its validation; ranking fidelity and savings for routine local evaluation remain unverified here.
Why it matters to Scott
The claimed ranking-preserving subset converges with Scott’s Progressive Evaluation Ladder: if validated, it offers a concrete cheaper screening step worth testing alongside his trace-backed agent comparisons. Ranking fidelity and savings remain unverified here, and fidelity within DeepSWE’s fixed harness would not establish transfer to Scott’s own harnesses or suitability as a regression quality gate.
ip:concept.progressive-evaluation-ladderip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisonradar:concept.agent-evaluationradar:concept.agent-benchmarks
queries asked of Scott's wikis
- coding-agent harness regression evaluation
- small benchmark subsets ranking fidelity
- evaluation cost versus statistical reliability
- model selection local coding-agent evaluations
- harness and reasoning effort leaderboard comparability
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 914h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p57 vs 519 stories at the 720h mark (now 914h old) — ahead of claude-subscriber-token-theft (1.0x), behind hydra-local-agentic-terminal (1.0x)
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-09-18T01:26:31Z
grounded: converges/medium — The claimed ranking-preserving subset converges with Scott’s Progressive Evaluation Ladder: if validated, it offers a concrete cheaper screening step worth test
2026-09-18T01:23:21Z
origin walked (codex/luna, conf 0.99): anchor reddit.post.1wjc5wj -> echo.other.56f46acf36 by codelion (published under the LocalLLaMA Hugging Face namespace)
2026-09-18T01:22:18Z
case created — A creator-announced downloadable dataset supplies a concrete, testable ranking-preservation claim directly relevant to inexpensive agent evaluation.
Decision trace
- 09-27 21:54review_dormantscheduled targets exhausted or 28 quiet days
- 09-27 21:54drop_targetsquiet through full ladder or over cap 8
- 09-19 13:22review_screenThe added comments express approval and restate the already-known relative-ranking rationale without new evidence or implementation results.
- 09-18 11:26groundThe claimed ranking-preserving subset converges with Scott’s Progressive Evaluation Ladder: if validated, it offers a concrete cheaper screening step worth testing alongside his trace-backed agent com
- 09-18 11:23promote_anchororigin walk conf 0.99
- 09-18 11:22createA creator-announced downloadable dataset supplies a concrete, testable ranking-preservation claim directly relevant to inexpensive agent evaluation.