2026-10-11 17:09 UTC

Independent evaluations will determine whether SWE-ContextBench reliably measures coding agents’ ability to learn repository-specific context and produces materially different findings from static software-engineering benchmarks.

state: expiredheat: lowuncertainty: highconvergesscott: highcoding-agents context-learning agent-evaluationSWE-ContextBench

What is this?

SWE-ContextBench is a repository-level benchmark designed to test whether coding agents can retrieve and reuse context from prior, related software-engineering tasks rather than merely solve isolated issues. The supplied paper snippets describe 1,100 base tasks and 376 related tasks spanning 51 repositories and nine programming languages; its authors report that accurately selected and summarized prior experience can improve resolution accuracy while reducing runtime and token costs, whereas irrelevant context can hurt. The snippets do not identify the people or organization behind SWE-ContextBench, and they do not substantiate the claim that independent evaluations have confirmed its reliability; some results also concern a separate, similarly named ContextBench benchmark.

Why it matters to Scott

SWE-ContextBench operationalizes Scott’s existing claims that coding-agent capability belongs to the model-plus-harness system, that repository learning must persist outside frozen models, and that selectively promoted context helps while irrelevant history can degrade performance. Its reported findings create a dated-receipts opportunity and could validate or challenge several load-bearing frameworks, although independent evaluation is still needed and the radar does not yet track this benchmark itself.
ip:source.a-blueprint-for-future-software-teamsip:concept.benchmarking-the-wrong-unitip:concept.model-plus-harness-benchmark-unitip:framework.context-engineeringip:concept.second-order-learningip:concept.memory-hygieneradar:concept.coding-agent-benchmarksradar:concept.agent-evaluationradar:concept.repository-intelligenceradar:memory-bench-layer-baseline-validity
queries asked of Scott's wikis
  • coding-agent memory across repository tasks
  • repository-specific context learning and reuse
  • coding-agent benchmark design beyond pass rates
  • retrieval quality versus context-window volume
  • agent experience summarization and negative transfer
  • coding harnesses with persistent project knowledge

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn ⭐Show HN: SWE-ContextBench – A Benchmark for Context Learning in Coding AgentsIreneAI20

Interpretation history

Decision trace