SWE-bench evaluates whether a system can resolve real GitHub issues by applying generated patches to fixed repository snapshots and running tests. Comparable scores require the same benchmark version and harness.
That last clause is essential. “SWE-bench score” is not one universal number unless the dataset, task subset, execution environment and scoring rules are the same.
On this page
What a SWE-bench task contains
The original 2024 paper introduced 2,294 problems drawn from 12 popular open-source Python repositories. Each task connects three things:
- a GitHub issue describing a bug or requested change;
- the repository at the base commit before the accepted fix;
- a merged pull request whose tests can show whether the issue was resolved.
The system does not receive only a short function-completion prompt. It must understand enough of a repository to locate relevant files, edit them consistently and produce a patch that runs in the project’s environment.
The paper reports an average repository size of hundreds of thousands of lines. That makes context selection part of the problem: a system has to find the small part of a large codebase that matters.
How tasks were selected
The benchmark authors began with pull requests from well-used Python repositories. Candidate tasks had to be merged, linked to an issue and include test changes.
They then applied execution-based filtering. A useful benchmark task needed at least one test whose result changed from failing before the accepted patch to passing after it. Tasks that could not be installed or executed reliably were removed.
This construction gives the task an observable outcome. It also means the benchmark represents the repositories, languages and maintenance practices in its dataset rather than every kind of software work.
How a generated patch is graded
A system starts from the repository’s base commit and the issue text. It returns a patch. The evaluation environment applies that patch and runs tests.
Two groups of tests matter conceptually:
- Fail-to-pass tests capture behaviour the accepted fix was meant to repair.
- Pass-to-pass tests check that existing behaviour still works.
A patch that makes the target test pass but breaks unrelated tests is not a clean resolution. This is why benchmark results are closer to repository repair than to a code-snippet quiz.
The benchmark does not directly measure product judgement, requirements discovery or whether a maintainer would accept the style of a change. It measures whether the submitted patch satisfies the benchmark’s executable checks.
Why retrieval and tools affect the result
The original paper compared sparse retrieval with an oracle setting. Sparse retrieval used issue text to select repository files. The oracle setting supplied files touched by the reference patch, which is useful for analysis but unrealistic as a normal workflow because it reveals where the accepted fix landed.
Modern coding systems may browse a repository, run commands, inspect test output and revise a patch across several steps. A benchmark report therefore needs to state the agent scaffold, tools, retry budget and compute limits, not only the underlying model.
A model inside one harness can produce a different result from the same model inside another. Treat the system under evaluation as model plus instructions, tools, environment and stopping rule.
What to compare in a result table
Before comparing two SWE-bench claims, check:
| Field | Why it matters |
|---|---|
| Benchmark variant | Full, Lite and verified subsets contain different tasks |
| Dataset version | Tasks and validation can change |
| Evaluation harness | Environment and patch application rules affect execution |
| Agent or scaffold | Retrieval, planning and tool use can change outcomes |
| Attempts or samples | Multiple attempts give a system more chances |
| Compute budget | Longer runs can inspect and revise more |
| Date | A current run should not be compared casually with an older setup |
A percentage without those fields is incomplete.
The LLM leaderboard keeps benchmark entries beside their stated source and date. Use the benchmark maintainer’s result or the lab’s technical report, then follow through to the methodology before treating a ranking as comparable.
What SWE-bench does not prove
A higher score is evidence that a configured system resolved more tasks in that benchmark run. It does not prove that the system:
- handles every programming language equally;
- understands a private company’s architecture;
- produces maintainable changes by default;
- works without human review;
- uses less time or money than a lower-scoring system.
Those questions need different evidence. Cost, latency, review burden and security boundaries can all matter more than a few percentage points for a real team.
A practical reading rule
Read a SWE-bench result in this order:
- Name the exact benchmark variant.
- Confirm the result source.
- Read the agent and tool setup.
- Check attempts and compute limits.
- Look for the evaluation harness and date.
- Compare only rows with materially similar conditions.
The headline score is the end of the method, not a substitute for it.




