Independent evaluations will determine whether SWE-ContextBench reliably measures coding agents’ ability to learn repository-specific context and produces materially different findings from static software-engineering benchmarks.
state: expiredheat: lowuncertainty: highconvergesscott: highcoding-agents context-learning agent-evaluationSWE-ContextBench
What is this?
SWE-ContextBench is a repository-level benchmark designed to test whether coding agents can retrieve and reuse context from prior, related software-engineering tasks rather than merely solve isolated issues. The supplied paper snippets describe 1,100 base tasks and 376 related tasks spanning 51 repositories and nine programming languages; its authors report that accurately selected and summarized prior experience can improve resolution accuracy while reducing runtime and token costs, whereas irrelevant context can hurt. The snippets do not identify the people or organization behind SWE-ContextBench, and they do not substantiate the claim that independent evaluations have confirmed its reliability; some results also concern a separate, similarly named ContextBench benchmark.
Why it matters to Scott
SWE-ContextBench operationalizes Scott’s existing claims that coding-agent capability belongs to the model-plus-harness system, that repository learning must persist outside frozen models, and that selectively promoted context helps while irrelevant history can degrade performance. Its reported findings create a dated-receipts opportunity and could validate or challenge several load-bearing frameworks, although independent evaluation is still needed and the radar does not yet track this benchmark itself.
ip:source.a-blueprint-for-future-software-teamsip:concept.benchmarking-the-wrong-unitip:concept.model-plus-harness-benchmark-unitip:framework.context-engineeringip:concept.second-order-learningip:concept.memory-hygieneradar:concept.coding-agent-benchmarksradar:concept.agent-evaluationradar:concept.repository-intelligenceradar:memory-bench-layer-baseline-validity
queries asked of Scott's wikis
- coding-agent memory across repository tasks
- repository-specific context learning and reuse
- coding-agent benchmark design beyond pass rates
- retrieval quality versus context-window volume
- agent experience summarization and negative transfer
- coding harnesses with persistent project knowledge
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-08-17T15:32:38Z
Repeated checks found no independent evaluation, implementation, adoption, or methodological discussion, and the original low-attention release has faded without establishing an active evaluation thread. The benchmark may be reconsidered if substantive third-party results emerge, but this episode no longer warrants routine monitoring.
2026-08-15T15:29:38Z
Repeated observation still finds no independent evaluation, adoption, or methodological scrutiny; the benchmark remains potentially relevant but unvalidated. Extend the watch interval rather than treating routine staleness as evidence for or against it.
2026-08-13T14:41:42Z
No independent evaluation, implementation, or methodological scrutiny has emerged, so the benchmark remains an unvalidated proposal rather than evidence that repository-context learning requires materially different evaluation.
2026-08-11T13:56:37Z
No independent evaluation, implementation, or methodological follow-up has appeared; the case remains a promising benchmark release whose reliability and differentiation from static benchmarks are untested.
2026-08-11T13:49:13Z
grounded: converges/high — SWE-ContextBench operationalizes Scott’s existing claims that coding-agent capability belongs to the model-plus-harness system, that repository learning must pe
2026-08-11T13:46:10Z
case created — The released benchmark targets a bounded coding-agent capability that existing static task evaluations may not measure.
Decision trace
- 08-18 01:32expireRepeated checks found no independent evaluation, implementation, adoption, or methodological discussion, and the original low-attention release has faded without establishing an active evaluation thre
- 08-18 01:32alert_silentThe only delta is elapsed time without new evidence; there is no consequential development to surface or near-term confirming fact to await.
- 08-18 01:32alert_routeThe only delta is elapsed time without new evidence; there is no consequential development to surface or near-term confirming fact to await.
- 08-16 01:29repriceRepeated observation still finds no independent evaluation, adoption, or methodological scrutiny; the benchmark remains potentially relevant but unvalidated. Extend the watch interval rather than trea
- 08-16 01:29alert_silentThere is no new consequential delta beyond elapsed time and unchanged engagement; wait for an independent reproduction, comparative result, or notable implementation.
- 08-16 01:29alert_routeThere is no new consequential delta beyond elapsed time and unchanged engagement; wait for an independent reproduction, comparative result, or notable implementation.
- 08-14 00:41repriceNo independent evaluation, implementation, or methodological scrutiny has emerged, so the benchmark remains an unvalidated proposal rather than evidence that repository-context learning requires mater
- 08-14 00:41alert_silentThe only trigger is elapsed staleness; there is no new consequential delta to surface. Revisit when an independent reproduction, comparative benchmark result, or notable implementation appears.
- 08-14 00:41alert_routeThe only trigger is elapsed staleness; there is no new consequential delta to surface. Revisit when an independent reproduction, comparative benchmark result, or notable implementation appears.
- 08-11 23:56repriceNo independent evaluation, implementation, or methodological follow-up has appeared; the case remains a promising benchmark release whose reliability and differentiation from static benchmarks are unt
- 08-11 23:56alert_silentThis is only an unchanged reobservation of the original low-engagement release, with no new consequential evidence; it can wait for independent results or adoption.
- 08-11 23:56alert_routeThis is only an unchanged reobservation of the original low-engagement release, with no new consequential evidence; it can wait for independent results or adoption.
- 08-11 23:51alert_silentThe Show HN establishes that a potentially relevant benchmark repository exists, but the supplied evidence contains no methodology, results, release details, or independent evaluation showing a conseq
- 08-11 23:51surface_candidateThe Show HN establishes that a potentially relevant benchmark repository exists, but the supplied evidence contains no methodology, results, release details, or independent evaluation showing a conseq
- 08-11 23:51alert_routeThe Show HN establishes that a potentially relevant benchmark repository exists, but the supplied evidence contains no methodology, results, release details, or independent evaluation showing a conseq
- 08-11 23:49groundSWE-ContextBench operationalizes Scott’s existing claims that coding-agent capability belongs to the model-plus-harness system, that repository learning must persist outside frozen models, and that se
- 08-11 23:46createThe released benchmark targets a bounded coding-agent capability that existing static task evaluations may not measure.