2026-10-11 16:37 UTC

Researchers from Meta Superintelligence Labs, Stanford, Harvard, and UW (SWE-bench lineage, led by Kilian Lieret and Ofir Press) released SWE-sweep β€” 100 repos, 4.1k bugs, where agents must find and fix unreported bugs with no hints β€” measuring proactive bug discovery at a stark 4.7% best (Sol 5.6 xhigh); leaderboard movement past that level or external adoption as a tracked agent-coding capability establishes proactive bug discovery as a benchmarked frontier, while stagnation marks the current gap as durable.

state: watchingheat: lowuncertainty: lowconvergesscott: mediumagent-evaluation coding-benchmarks proactive-bug-discovery agent-harnessesKilian LieretOfir PressJohn YangMeta Superintelligence Labs

What is this?

Per the case's own evidence (a Show HN post), researchers in the SWE-bench lineage β€” Kilian Lieret and Ofir Press's circle, spanning Meta Superintelligence Labs, Stanford, Harvard, and UW β€” released SWE-sweep: 100 real repositories and ~4.1k bugs where coding agents must find and fix defects nobody reported, with the best configuration (Sol 5.6 xhigh) scoring only 4.7%. The supplied web results do not surface SWE-sweep itself, so its specific figures rest on the case's evidence; the snippets do corroborate the lineage (Lieret, John Yang, and Press appear in SWE-agent and SWE-smith author lists). What the results do show is that 'proactive bug fixing without issue reports' became a crowded 2026 benchmark frontier from multiple independent teams: Active-SWE (1,663 tasks, six bug categories, dual-track evaluation of recorded-bug resolution at ~20% best plus potential-bug discovery, with Claude Opus 4.8 and GLM-5.2 leading) alongside SWE-atlas, ChainSWE, ProgramBench (0–0.5% best), and SWE-bench Science (47.9% best) β€” all measuring agents 'beyond issue resolution' and all finding steep capability gaps. That parallel-artifact wave partially corroborates the case's 'adoption' arm while meaning the resolvable signal is leaderboard movement within this now-busy space, not the concept's debut.

Why it matters to Scott

The benchmark's headline β€” agents fix only ~4.7% of bugs nobody reported β€” is a dated, credible-team receipt for the exact premise Scott's audit architectures are built on: his security-reviewer method deliberately targets closed checkable attacker paths and bounded-security-context-escalation exists because unguided discovery fails, so the SWE-bench lineage has now publicly quantified what his canon argues rather than challenging it. The watchable value is calibration and timing: this leaderboard tells him when heartbeat-style audit charters can start delegating real discovery authority, though per the grounding the space is already crowding (Active-SWE scores potential-bug discovery too), so the live signal is leaderboard movement here specifically, not the concept's debut.
ip:concept.unsolicited-corroborationip:framework.heartbeat-supervisory-programip:source.security-reviewer-method-ebookdev:concept.bounded-security-context-escalationdev:project.wordpress-security-reviewradar:concept.coding-agent-benchmarksradar:concept.agent-evaluationradar:lovable-hacking-agent-swarmsradar:cloudflare-security-audit-skillradar:harnesseval-code-review-gainsradar:specific-real-swe-releaseradar:post-merge-agentic-code-benchmark
queries asked of Scott's wikis
  • harness engineering vs model upgrades coding agents
  • proactive vs reactive agent workflow code audit
  • SWE-bench benchmark saturation critique
  • agent code review pipeline static analysis verification
  • agent autonomy unwanted changes guardrails
  • coding agent evaluation design reliability

Measured heat

now 0 pts/hpeak 31 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 266h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-30 14:00⭐ origin echo-reconstructedOriginal announcement of the SWE-sweep benchmark: "How many bugs can LMs find & fix in large codebases? ... Given a codebase containing many
Kilian Lieret, Jeffrey Jian Ma, Rahul Kindi, Yuxiang Wei, Jeremy Ma, Sten Sootla, Parth Thakkar, Chao Beyond Zhou, Pengcheng Yin, Rui Hou, Ofir Press, John Yang β€” Meta Superintelligence Labs, Harvard, UW, Stanford (site Β© Meta Platforms, Inc) on other (echo) Β· attributed from hn.story.49923102
β€”
10-01 15:36first on hacker news Β· published Β· +25.6hShow HN: Benchmark: AI doesn't find bugs unless you tell it what's wrong
lieret
β€”
10-02 15:55first on r/OpenAI Β· published Β· +49.9hLuna and Sol doing extremely well on new benchmark about finding bugs before users run into them
klieret
β€”
10-02 16:05first on r/LocalLLaMA Β· published Β· +50.1hNew benchmark on LMs fixing bugs before users run into them
klieret
β€”
10-01 15:36amplified on hacker newshn.story.49923102
lieret
peak 5 Β· 2 comments Β· 25% of case engagement
10-02 15:55amplified on r/OpenAIreddit.post.1wvxg2p
klieret
peak 9 Β· 2 comments Β· 22% of case engagement
10-02 16:05amplified on r/LocalLLaMA πŸ‘‘reddit.post.1wvxph8
klieret
peak 12 Β· 14 comments Β· 52% of case engagement
10-01 16:21our radar first saw it Β· +26.4hdiscovery anchor: hn.story.49923102β€”
pace: p56 vs 1188 stories at the 168h mark (now 266h old) β€” ahead of chatgpt-word-integration (1.0x), behind opencontext-project-local-agent-memory (1.0x)

Evidence (4) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Benchmark: AI doesn't find bugs unless you tell it what's wrong
Retrieved article excerpt

Open article Β· Retrieved 2026-10-01T16:32:07.352488+00:00

Capybara sweeping bugs
Capybara sweeping bugs

# SWE-sweep

How many bugs can LMs find & fix in large codebases?

Given a real repository, an agent must discover & repair as many bugs as they can. Agents are not given any hint about the type of bug or its location.

[Kilian Lieret1](https://www.lieret.net/) Β· [Jeffrey Jian Ma1,2](https://18jeffreyma.github.io/) Β· [Rahul Kindi1](https://github.com/rkindi) Β· [Yuxiang Wei1](https://yuxiang.cs.illinois.edu)

[Jeremy Ma1,3](https://github.com/Awayfaring) Β· [Sten Sootla1](https://scholar.google.com/citations?user=UAx_woYAAAAJ&hl=en) Β· [Parth Thakkar1](https://thakkarparth007.github.io/) Β· [Chao Beyond Zhou1](https://github.com/think-step-by-step)

[Pengcheng Yin1](https://pengcheng.in/) Β· [Rui Hou1](https://scholar.google.com/citations?user=PKHKqX0AAAAJ&hl=en) Β· [Ofir Press1](https://ofir.io/) Β· [John Yang1,4](https://john-b-yang.github.io/)

[1 Meta Superintelligence Labs](https://ai.meta.com/) Β· [2 Harvard University](https://g.harvard.edu) Β· [3 University of Washington](https://uw.edu) Β· [4 Stanford University](https://stanford.edu)

100 repositories Β· 4.1k bugs Β· Updated September 24, 2026

LeaderboardDetailsPareto

| Rank |  | Model | Agent | Bugs resolvedScore | Total USDUSD | USD / repo | Turns / repo | Tokens / repo |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | OpenAI | Sol 5.6 (xhigh)OpenAI | mini-SWE-agent | 4.7% | $7,230 | $72.30 | 233 | 104.8k |
| 2 | OpenAI | Luna 5.6 (xhigh)OpenAI | mini-SWE-agent | 2.5% | $224 | $2.24 | 204 | 61.2k |
| 3 | OpenAI | Terra 5.6 (xhigh)OpenAI | mini-SWE-agent | 1.5% | $357 | $3.57 | 76 | 46.6k |
| 4 | OpenAI | Luna 5.6 (high)OpenAI | mini-SWE-agent | 1.4% | $28 | $0.28 | 75 | 20.3k |
| 5 | Anthropic | Opus 5 (xhigh)Anthropic | mini-SWE-agent | 1.3% | $5,363 | $53.63 | 323 | 192.7k |
| 6 | Moonshot AI | Kimi K3Moonshot AI | mini-SWE-agent | 0.6% | $2,451 | $24.51 | 337 | 152.6k |
| 7 | OpenAI | Luna 5.6OpenAI | mini-SWE-agent | 0.5% | $4 | $0.04 | 22 | 4.5k |
| 8 | OpenAI | GPT-5.4 Mini (high)OpenAI | mini-SWE-agent | 0.5% | $122 | $1.22 | 75 | 40.5k |
| 9 | OpenAI | GPT-5.4 MiniOpenAI | mini-SWE-agent | 0.2% | $5 | $0.05 | 13 | 2.1k |
| 10 | Google | Gemini 3.5 Flash LiteGoogle | mini-SWE-agent | 0.1% | $6 | $0.06 | 34 | 5.2k |

Per-repository resource use includes non-deprecated retries and is averaged over evaluated repositories.

Resource use is averaged per evaluated repository

LeaderboardDetailsPareto

CostTurnsTokens

Model
Resolved / Cost

[1
OpenAI

Sol 5.6 (xhigh)
OpenAI

4.7%
$72.30](https://swesweep.com/model/sol-5-6-xhigh/)
[2
OpenAI

Luna 5.6 (xhigh)
OpenAI

2.5%
$2.24](https://swesweep.com/model/luna-5-6-xhigh/)
[3
OpenAI

Terra 5.6 (xhigh)
OpenAI

1.5%
$3.57](https://swesweep.com/model/terra-5-6-xhigh/)
[4
OpenAI

Luna 5.6 (high)
OpenAI

1.4%
$0.28](https://swesweep.com/model/luna-5-6-high/)
[5
Anthropic

Opus 5 (xhigh)
Anthropic

1.3%
$53.63](https://swesweep.com/model/opus-5-xhigh/)
[6
Moonshot AI

Kimi K3
Moonshot AI

0.6%
$24.51](https://swesweep.com/model/kimi-k3/)
[7
OpenAI

Luna 5.6
OpenAI

0.5%
$0.04](https://swesweep.com/model/luna-5-6/)
[8
OpenAI

GPT-5.4 Mini (high)
OpenAI

0.5%
$1.22](https://swesweep.com/model/gpt-5-4-mini-high/)
[9
OpenAI

GPT-5.4 Mini
OpenAI

0.2%
$0.05](https://swesweep.com/model/gpt-5-4-mini/)
[10
Google

Gemini 3.5 Flash Lite
Google

0.1%
$0.06](https://swesweep.com/model/gemini-3-5-flash-lite/)

Hover a point for details Β· The line marks the Pareto frontier (best result per cost) Β· Click a point to see model details

## About

Most existing software engineering benchmarks evaluate coding agents on concrete, well-specified tasks, commonly by providing a codebase together with a user-reported issue to resolve. However, as users delegate increasingly broad outcomes to coding agents, the natural next step is for agents to determine not only how to perform useful work, but also what useful work needs to be done.

An agent entrusted with a repository should be able to decide what is broken, which problems matter, and how to solve them before they are reported. We introduce SWE-sweep, a benchmark for this open-ended setting.

An agent explores a codebase, finds multiple bugs across its files, and outputs a patch to fix them.

Given a codebase containing many concurrent bugs and no information about their nature or location, an agent must autonomously discover and fix as many bugs as possible.

SWE-sweep is constructed from open-source repositories. For each repository, we collect issue-pull request pairs, then identify a single commit where the maximum number of bugs are present at the same time.

A repository timeline showing how a snapshot with four concurrently present bugs is selected.

Each repair is [evaluated](https://swesweep.com/#faq-evaluation) against hidden tests from the corresponding pull requests, along with the existing test suite to check for regressions.

Success requires agents to explore and understand a large codebase over long horizon work, repair bugs without introducing regressions, and manage interactions among fixes that are not independent.

How are tasks constructed?

We collect real issue–pull request pairs, identify a commit where many of those bugs coexist, and retain bugs whose fixes and tests can be reproduced at that repository state.
Besides many quality filters shared with other benchmarks, we apply [extensive filtering](https://swesweep.com/#faq-bug-selection) to evaluate only bugs that can be [discovered from reading the repository alone](https://swesweep.com/#faq-discoverable).

What does an agent receive?

A repository at a fixed base commit and a broad instruction to find and fix as many bugs as possible. It receives no issue descriptions, filenames, line ranges, or other bug-specific hints.

You can find the [full prompt here](https://swesweep.com/prompt/).

How is SWE-sweep evaluated?

We score every task against a reference set of previously identified bugs (see [Construction](https://swesweep.com/#faq-construction), counting how many the agent successfully repairs.

For every task, we run the agent's submitted codebase against two sets of tests.
First, we restore the repository's original test suite and run it to verify no existing behavior was broken.
Second, for each bug, we run a set of hidden tests; at least one of these tests fails on the unmodified codebase, and passes once the bug is fixed (fail-to-pass).
A bug is considered resolved if all its hidden tests pass and the original suite still passes.

The benchmark score is the fraction of all bugs across all repositories that have been resolved.
[What about any other changes that the agent makes?](https://swesweep.com/#faq-other-changes)

What bugs are in the benchmark? How do you guarantee the task is feasible?

We filter bugs (represented by a test patch and a fix patch) to ensure the task is feasible. The criteria are:

1. The test patch and fix patch independently apply to the base commit.
2. All target tests pass after the fix patch has been applied to the base commit, but at least one target test fails on the base commit (*F2P tests*). There might be additional tests that pass before and after the fix patch has been applied (*P2P tests*).
3. The F2P tests reveal a single [*discoverable bug*](https://swesweep.com/#faq-discoverable) in the base commit.
4. The target tests are not overly specific; any reasonable fix to the discoverable bug will pass the target tests.
5. The target tests do not contradict the base commit tests.
6. Target tests of different bug instances do not contradict each other.

Appendix A.2 in the [paper](https://swesweep.com/paper) discusses feasibility in detail.

What makes a bug discoverable?

The expected behavior must be inferable from the repository itself, for example through documentation, types, existing tests, callers, invariants, standards, or an unambiguously undesirable failure such as a crash or data loss.
The latter category is used extremely conservatively and all but 2 bugs in the benchmark have concrete repository contracts that describe the expected behavior.
You can find some examples about what we mean with *repository contracts* [here](https://swesweep.com/discoverability/).
We have spent a lot of time validating this aspect of the benchmark and you can find more details in the appendix of our [paper](https://swesweep.com/paper).

What about any other changes that the agent makes?

Any history-derived benchmark necessarily under-counts the bugs present in a repository.
By restricting PandoraBench to defects confirmed by an upstream fix, we ensure that every bug in the benchmark is backed by strong evidence that the observed behavior was considered erroneous by the repository maintainers.
We therefore only score the agent's changes on the bugs that are confirmed by an upstream fix and supported by executable regression tests, as well as the other [quality filters](https://swesweep.com/#faq-bug-selection).
However, if an agent causes a regression in the original test suite (the agent is [explicitly told](https://swesweep.com/prompt/) to avoid this), it will be scores as 0%.
This means that the agent's changes that are not scored by the set of bugs are still likely to be non-destructive and compatible with the repository's existing behavior.
See A.3 and A.4 in the [paper](https://swesweep.com/#faq-other-changes) for more discussion.

Does more inference-time compute help?

Repeated attempts recover additional bugs, but the gains diminish. Later work within one run can also undo earlier repairs, so simply extending a trajectory does not guarantee improvement.
See Fig. 7 in the [paper](https://swesweep.com/paper).

What about Astra, Fable, 5.5, ...?

We're working on evaluating more models! The current selection was finalized for our ICLR submission. We're also looking into even higher reasoning modes, but this might push over $10k for a single run. We also want to have more open weights models on the leaderboard.

Why mini-swe-agent? Could other scaffolds/multiagents achieve higher performance?

Our paper has an ablation with Claude Code and Codex. Neither seems to significantly outperform mini-swe-agent (to the contrary, mini-swe-agent is even quite a bit better than Codex). This follows many other benchmarks, where mini-swe-agent has been extremely competitive. However, we absolutely hope to kick off more research into the role of agent scaffolds and will [open for submissions soon](https://swesweep.com/#faq-submit).

How do I submit to the leaderboard?

Public submissions are coming soon.

[Browse repositories
Explore all 100 benchmark repositories and their results.](https://swesweep.com/repositories/)

## Citation

```
@misc{lieret2026swesweep,
  title  = {{SWE-sweep}: Can Agents Autonomously Find and Fix Bugs?},
  author = {Kilian Lieret and Jeffrey Jian Ma and Rahul Kindi and
            Yuxiang Wei and Jeremy Ma and Sten Sootla and
            Parth Thakkar and Chao Beyond Zhou and Pengcheng Yin and
            Rui Hou and Ofir Press and John Yang},
  year   = {2026},
  note   = {Preprint}
}
```
lieret52
🟧 echo.other ⭐Original announcement of the SWE-sweep benchmark: "How many bugs can LMs find & fix in large codebases? ... Given a codebase containing manyKilian Lieret, Jeffrey Jian Ma, Rahul Kindi, Yuxiang Wei, Jeremy Ma, Sten Sootla, Parth Thakkar, Chao Beyond Zhou, Pengcheng Yin, Rui Hou, Ofir Press, John Yang β€” Meta Superintelligence Labs, Harvard, UW, Stanford (site Β© Meta Platforms, Inc)β€”β€”
🟠 redditNew benchmark on LMs fixing bugs before users run into them
LocalLLaMA
klieret1214
🟠 redditLuna and Sol doing extremely well on new benchmark about finding bugs before users run into them
OpenAI
klieret92

Interpretation history

Decision trace