2026-10-11 17:10 UTC

DeepSWE-mini creator asankhs claims the released 16-instance subset preserves the full DeepSWE leaderboard's relative model rankings, potentially reducing the cost of routine local coding-agent evaluation without reproducing absolute scores.

state: seedheat: mediumuncertainty: mediumconvergesscott: mediumagent-evaluation coding-agents benchmarkingasankhs

What is this?

DeepSWE is Datacurve’s benchmark for long-horizon coding agents, covering 113 original tasks across 91 repositories and five languages, with isolated environments, executable verifiers, and a fixed mini-swe-agent harness. The case attributes DeepSWE-mini to asankhs, who claims a released 16-task subset preserves the full leaderboard’s relative rankings rather than its absolute scores. The supplied web snippets establish the parent benchmark but do not directly document the mini release, its creator, or its validation; ranking fidelity and savings for routine local evaluation remain unverified here.

Why it matters to Scott

The claimed ranking-preserving subset converges with Scott’s Progressive Evaluation Ladder: if validated, it offers a concrete cheaper screening step worth testing alongside his trace-backed agent comparisons. Ranking fidelity and savings remain unverified here, and fidelity within DeepSWE’s fixed harness would not establish transfer to Scott’s own harnesses or suitability as a regression quality gate.
ip:concept.progressive-evaluation-ladderip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisonradar:concept.agent-evaluationradar:concept.agent-benchmarks
queries asked of Scott's wikis
  • coding-agent harness regression evaluation
  • small benchmark subsets ranking fidelity
  • evaluation cost versus statistical reliability
  • model selection local coding-agent evaluations
  • harness and reasoning effort leaderboard comparability

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 914h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-03 14:00⭐ origin echo-reconstructedThe primary artifact is the dataset commit titled “deepswe-mini at 16 tasks, three-way validated.” Its README says it contains “16 of the 11
codelion (published under the LocalLLaMA Hugging Face namespace) on other (echo) · attributed from reddit.post.1wjc5wj
—
09-18 01:13first on r/LocalLLaMA · published · +347.2hDeepSWE-mini a 16 instance subset of DeepSWE that replicates the rankings of the leaderboard
asankhs
—
09-18 01:13amplified on r/LocalLLaMA 👑reddit.post.1wjc5wj
asankhs
peak 25 · 6 comments · 100% of case engagement
09-18 01:20our radar first saw it · +347.3hdiscovery anchor: reddit.post.1wjc5wj—
pace: p57 vs 519 stories at the 720h mark (now 914h old) — ahead of claude-subscriber-token-theft (1.0x), behind hydra-local-agentic-terminal (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditDeepSWE-mini a 16 instance subset of DeepSWE that replicates the rankings of the leaderboard
LocalLLaMA
asankhs256
🟧 echo.other ⭐The primary artifact is the dataset commit titled “deepswe-mini at 16 tasks, three-way validated.” Its README says it contains “16 of the 11codelion (published under the LocalLLaMA Hugging Face namespace)——

Interpretation history

Decision trace