Retrieved article excerpt
Open article Β· Retrieved 2026-10-01T16:32:07.352488+00:00
Capybara sweeping bugs
Capybara sweeping bugs
# SWE-sweep
How many bugs can LMs find & fix in large codebases?
Given a real repository, an agent must discover & repair as many bugs as they can. Agents are not given any hint about the type of bug or its location.
[Kilian Lieret1](https://www.lieret.net/) Β· [Jeffrey Jian Ma1,2](https://18jeffreyma.github.io/) Β· [Rahul Kindi1](https://github.com/rkindi) Β· [Yuxiang Wei1](https://yuxiang.cs.illinois.edu)
[Jeremy Ma1,3](https://github.com/Awayfaring) Β· [Sten Sootla1](https://scholar.google.com/citations?user=UAx_woYAAAAJ&hl=en) Β· [Parth Thakkar1](https://thakkarparth007.github.io/) Β· [Chao Beyond Zhou1](https://github.com/think-step-by-step)
[Pengcheng Yin1](https://pengcheng.in/) Β· [Rui Hou1](https://scholar.google.com/citations?user=PKHKqX0AAAAJ&hl=en) Β· [Ofir Press1](https://ofir.io/) Β· [John Yang1,4](https://john-b-yang.github.io/)
[1 Meta Superintelligence Labs](https://ai.meta.com/) Β· [2 Harvard University](https://g.harvard.edu) Β· [3 University of Washington](https://uw.edu) Β· [4 Stanford University](https://stanford.edu)
100 repositories Β· 4.1k bugs Β· Updated September 24, 2026
LeaderboardDetailsPareto
| Rank | | Model | Agent | Bugs resolvedScore | Total USDUSD | USD / repo | Turns / repo | Tokens / repo |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | OpenAI | Sol 5.6 (xhigh)OpenAI | mini-SWE-agent | 4.7% | $7,230 | $72.30 | 233 | 104.8k |
| 2 | OpenAI | Luna 5.6 (xhigh)OpenAI | mini-SWE-agent | 2.5% | $224 | $2.24 | 204 | 61.2k |
| 3 | OpenAI | Terra 5.6 (xhigh)OpenAI | mini-SWE-agent | 1.5% | $357 | $3.57 | 76 | 46.6k |
| 4 | OpenAI | Luna 5.6 (high)OpenAI | mini-SWE-agent | 1.4% | $28 | $0.28 | 75 | 20.3k |
| 5 | Anthropic | Opus 5 (xhigh)Anthropic | mini-SWE-agent | 1.3% | $5,363 | $53.63 | 323 | 192.7k |
| 6 | Moonshot AI | Kimi K3Moonshot AI | mini-SWE-agent | 0.6% | $2,451 | $24.51 | 337 | 152.6k |
| 7 | OpenAI | Luna 5.6OpenAI | mini-SWE-agent | 0.5% | $4 | $0.04 | 22 | 4.5k |
| 8 | OpenAI | GPT-5.4 Mini (high)OpenAI | mini-SWE-agent | 0.5% | $122 | $1.22 | 75 | 40.5k |
| 9 | OpenAI | GPT-5.4 MiniOpenAI | mini-SWE-agent | 0.2% | $5 | $0.05 | 13 | 2.1k |
| 10 | Google | Gemini 3.5 Flash LiteGoogle | mini-SWE-agent | 0.1% | $6 | $0.06 | 34 | 5.2k |
Per-repository resource use includes non-deprecated retries and is averaged over evaluated repositories.
Resource use is averaged per evaluated repository
LeaderboardDetailsPareto
CostTurnsTokens
Model
Resolved / Cost
[1
OpenAI
Sol 5.6 (xhigh)
OpenAI
4.7%
$72.30](https://swesweep.com/model/sol-5-6-xhigh/)
[2
OpenAI
Luna 5.6 (xhigh)
OpenAI
2.5%
$2.24](https://swesweep.com/model/luna-5-6-xhigh/)
[3
OpenAI
Terra 5.6 (xhigh)
OpenAI
1.5%
$3.57](https://swesweep.com/model/terra-5-6-xhigh/)
[4
OpenAI
Luna 5.6 (high)
OpenAI
1.4%
$0.28](https://swesweep.com/model/luna-5-6-high/)
[5
Anthropic
Opus 5 (xhigh)
Anthropic
1.3%
$53.63](https://swesweep.com/model/opus-5-xhigh/)
[6
Moonshot AI
Kimi K3
Moonshot AI
0.6%
$24.51](https://swesweep.com/model/kimi-k3/)
[7
OpenAI
Luna 5.6
OpenAI
0.5%
$0.04](https://swesweep.com/model/luna-5-6/)
[8
OpenAI
GPT-5.4 Mini (high)
OpenAI
0.5%
$1.22](https://swesweep.com/model/gpt-5-4-mini-high/)
[9
OpenAI
GPT-5.4 Mini
OpenAI
0.2%
$0.05](https://swesweep.com/model/gpt-5-4-mini/)
[10
Google
Gemini 3.5 Flash Lite
Google
0.1%
$0.06](https://swesweep.com/model/gemini-3-5-flash-lite/)
Hover a point for details Β· The line marks the Pareto frontier (best result per cost) Β· Click a point to see model details
## About
Most existing software engineering benchmarks evaluate coding agents on concrete, well-specified tasks, commonly by providing a codebase together with a user-reported issue to resolve. However, as users delegate increasingly broad outcomes to coding agents, the natural next step is for agents to determine not only how to perform useful work, but also what useful work needs to be done.
An agent entrusted with a repository should be able to decide what is broken, which problems matter, and how to solve them before they are reported. We introduce SWE-sweep, a benchmark for this open-ended setting.
An agent explores a codebase, finds multiple bugs across its files, and outputs a patch to fix them.
Given a codebase containing many concurrent bugs and no information about their nature or location, an agent must autonomously discover and fix as many bugs as possible.
SWE-sweep is constructed from open-source repositories. For each repository, we collect issue-pull request pairs, then identify a single commit where the maximum number of bugs are present at the same time.
A repository timeline showing how a snapshot with four concurrently present bugs is selected.
Each repair is [evaluated](https://swesweep.com/#faq-evaluation) against hidden tests from the corresponding pull requests, along with the existing test suite to check for regressions.
Success requires agents to explore and understand a large codebase over long horizon work, repair bugs without introducing regressions, and manage interactions among fixes that are not independent.
How are tasks constructed?
We collect real issueβpull request pairs, identify a commit where many of those bugs coexist, and retain bugs whose fixes and tests can be reproduced at that repository state.
Besides many quality filters shared with other benchmarks, we apply [extensive filtering](https://swesweep.com/#faq-bug-selection) to evaluate only bugs that can be [discovered from reading the repository alone](https://swesweep.com/#faq-discoverable).
What does an agent receive?
A repository at a fixed base commit and a broad instruction to find and fix as many bugs as possible. It receives no issue descriptions, filenames, line ranges, or other bug-specific hints.
You can find the [full prompt here](https://swesweep.com/prompt/).
How is SWE-sweep evaluated?
We score every task against a reference set of previously identified bugs (see [Construction](https://swesweep.com/#faq-construction), counting how many the agent successfully repairs.
For every task, we run the agent's submitted codebase against two sets of tests.
First, we restore the repository's original test suite and run it to verify no existing behavior was broken.
Second, for each bug, we run a set of hidden tests; at least one of these tests fails on the unmodified codebase, and passes once the bug is fixed (fail-to-pass).
A bug is considered resolved if all its hidden tests pass and the original suite still passes.
The benchmark score is the fraction of all bugs across all repositories that have been resolved.
[What about any other changes that the agent makes?](https://swesweep.com/#faq-other-changes)
What bugs are in the benchmark? How do you guarantee the task is feasible?
We filter bugs (represented by a test patch and a fix patch) to ensure the task is feasible. The criteria are:
1. The test patch and fix patch independently apply to the base commit.
2. All target tests pass after the fix patch has been applied to the base commit, but at least one target test fails on the base commit (*F2P tests*). There might be additional tests that pass before and after the fix patch has been applied (*P2P tests*).
3. The F2P tests reveal a single [*discoverable bug*](https://swesweep.com/#faq-discoverable) in the base commit.
4. The target tests are not overly specific; any reasonable fix to the discoverable bug will pass the target tests.
5. The target tests do not contradict the base commit tests.
6. Target tests of different bug instances do not contradict each other.
Appendix A.2 in the [paper](https://swesweep.com/paper) discusses feasibility in detail.
What makes a bug discoverable?
The expected behavior must be inferable from the repository itself, for example through documentation, types, existing tests, callers, invariants, standards, or an unambiguously undesirable failure such as a crash or data loss.
The latter category is used extremely conservatively and all but 2 bugs in the benchmark have concrete repository contracts that describe the expected behavior.
You can find some examples about what we mean with *repository contracts* [here](https://swesweep.com/discoverability/).
We have spent a lot of time validating this aspect of the benchmark and you can find more details in the appendix of our [paper](https://swesweep.com/paper).
What about any other changes that the agent makes?
Any history-derived benchmark necessarily under-counts the bugs present in a repository.
By restricting PandoraBench to defects confirmed by an upstream fix, we ensure that every bug in the benchmark is backed by strong evidence that the observed behavior was considered erroneous by the repository maintainers.
We therefore only score the agent's changes on the bugs that are confirmed by an upstream fix and supported by executable regression tests, as well as the other [quality filters](https://swesweep.com/#faq-bug-selection).
However, if an agent causes a regression in the original test suite (the agent is [explicitly told](https://swesweep.com/prompt/) to avoid this), it will be scores as 0%.
This means that the agent's changes that are not scored by the set of bugs are still likely to be non-destructive and compatible with the repository's existing behavior.
See A.3 and A.4 in the [paper](https://swesweep.com/#faq-other-changes) for more discussion.
Does more inference-time compute help?
Repeated attempts recover additional bugs, but the gains diminish. Later work within one run can also undo earlier repairs, so simply extending a trajectory does not guarantee improvement.
See Fig. 7 in the [paper](https://swesweep.com/paper).
What about Astra, Fable, 5.5, ...?
We're working on evaluating more models! The current selection was finalized for our ICLR submission. We're also looking into even higher reasoning modes, but this might push over $10k for a single run. We also want to have more open weights models on the leaderboard.
Why mini-swe-agent? Could other scaffolds/multiagents achieve higher performance?
Our paper has an ablation with Claude Code and Codex. Neither seems to significantly outperform mini-swe-agent (to the contrary, mini-swe-agent is even quite a bit better than Codex). This follows many other benchmarks, where mini-swe-agent has been extremely competitive. However, we absolutely hope to kick off more research into the role of agent scaffolds and will [open for submissions soon](https://swesweep.com/#faq-submit).
How do I submit to the leaderboard?
Public submissions are coming soon.
[Browse repositories
Explore all 100 benchmark repositories and their results.](https://swesweep.com/repositories/)
## Citation
```
@misc{lieret2026swesweep,
title = {{SWE-sweep}: Can Agents Autonomously Find and Fix Bugs?},
author = {Kilian Lieret and Jeffrey Jian Ma and Rahul Kindi and
Yuxiang Wei and Jeremy Ma and Sten Sootla and
Parth Thakkar and Chao Beyond Zhou and Pengcheng Yin and
Rui Hou and Ofir Press and John Yang},
year = {2026},
note = {Preprint}
}
```