Edge Rewrite
Jump to content

SWE-bench

From Wikipedia, the free encyclopedia

SWE-bench is a benchmark for large language models and coding agents. Its 2,294 tasks were taken from issues that were actually filed, and actually fixed, in 12 Python repositories: astropy, django, flask, matplotlib, pylint, pytest, requests, scikit-learn, seaborn, sphinx, sympy and xarray. Django alone supplies 850 of them, sympy another 386. The system being tested is shown the repository as it stood before the fix, plus the text of the issue, and has to produce a patch. It never sees the tests that will judge the result. When the benchmark was introduced in October 2023 by a group at Princeton University and the University of Chicago, under the title "Can Language Models Resolve Real-World GitHub Issues?", the strongest model they tried, Claude 2, fixed only 1.96% of the tasks.[1]

In August 2024, OpenAI released SWE-bench Verified, a human-validated subset of SWE-bench. In February 2026, the leading published result on SWE-bench Verified was 80.9%. In a post titled "Why SWE-bench Verified no longer measures frontier coding capabilities", OpenAI said it would stop quoting the figure at all. The models being scored had read the repositories the tasks come from, and in the sample OpenAI examined every frontier model it tried could reproduce wording from the problem statements or from the developers' own fix. OpenAI recommended switching to SWE-bench Pro, another variant.[2]

Variants

[edit]

SWE-bench Verified, put together with OpenAI in August 2024, is the version model releases usually quote: 500 tasks whose descriptions and tests were read through by human annotators first.[3][2] Three spin-offs followed: Lite (300 tasks that each touch a single file, cheaper to run), Multimodal (617 tasks from JavaScript interface libraries, each with an image in the issue or the tests) and Multilingual (300 tasks in nine languages).[4][5][6] Scale AI's SWE-Bench Pro, from September 2025, answers the contamination problem by keeping much of itself secret: of its 1,865 tasks, one part is withheld and another comes from private repositories.[7] A 2026 paper found the same leakage there and published a corrected version.[8]

Criticism

[edit]

Reem Aleithan and colleagues read through the patches credited in 2024 to SWE-agent with GPT-4, which was then top of the leaderboard. In 32.67% of the tasks it had supposedly solved, the fix was written out in the issue report or in the comments under it. Another 31.08% passed only because the tests were too weak to catch a patch that did the wrong thing. Take both groups out and the score drops from 12.47% to 3.97%. More than 94% of the issues had been filed before the models being tested finished training.[9]

A second study, asking whether the "solved issues" in SWE-bench were really solved correctly, compared accepted patches against the fixes the projects' own developers had written. Of the patches the benchmark counts as correct, 7.8% fail the developers' test suite outright, and 29.6% of the rest behave differently from the real fix, usually by changing more of the program than the developers changed.[10] The tests cut the other way as well: OpenAI, which had helped build Verified, found that at least 59.4% of the problems it sampled carry tests that reject a patch even when the patch works. It told other developers to stop quoting Verified and to use SWE-Bench Pro instead.[2]

Who was doing the submitting changed as well. A 2026 survey of the Lite and Verified leaderboards, 79 entries on one and 133 on the other, found most of the work coming from industry rather than universities, and proprietary models behind almost every leading entry.[11]

See also

[edit]

References

[edit]
  1. ↑ Jimenez, Carlos E.; Yang, John; Wettig, Alexander; Yao, Shunyu; Pei, Kexin; Press, Ofir; Narasimhan, Karthik (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. Twelfth International Conference on Learning Representations. arXiv:2310.06770.
  2. 1 2 3 "Why SWE-bench Verified no longer measures frontier coding capabilities". OpenAI. 23 February 2026. Retrieved 1 October 2026.
  3. ↑ "SWE-bench Verified". swebench.com. Retrieved 1 October 2026.
  4. ↑ "SWE-bench Lite". swebench.com. Retrieved 1 October 2026.
  5. ↑ Yang, John; Jimenez, Carlos E.; Zhang, Alex L.; et al. (4 October 2024). "SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?". arXiv:2410.03859 [cs.SE].
  6. ↑ Khandpur, Kabir; Lieret, Kilian; Jimenez, Carlos E.; et al. "SWE-bench Multilingual". swebench.com. Retrieved 1 October 2026.
  7. ↑ Deng, Xiang; Da, Jeff; Pan, Edwin; et al. (21 September 2025). "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?". arXiv:2509.16941 [cs.SE].
  8. ↑ Zheng, Pujun; Shang, Zixin; Jiang, Shufan; et al. (8 September 2026). "SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents". arXiv:2609.08149 [cs.SE].
  9. ↑ Aleithan, Reem; Xue, Haoran; Mohajer, Mohammad Mahdi; Nnorom, Elijah; Uddin, Gias; Wang, Song (9 October 2024). "SWE-Bench+: Enhanced Coding Benchmark for LLMs". arXiv:2410.06992 [cs.SE].
  10. ↑ Wang, You; Pradel, Michael; Liu, Zhongxin (2026). Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study. 2026 IEEE/ACM 48th International Conference on Software Engineering. pp. 169–181. arXiv:2503.15223. doi:10.1145/3744916.3764576.
  11. ↑ Martinez, Matias; Franch, Xavier (2026). What's in a Benchmark? The Case of SWE-Bench in Automated Program Repair. IEEE/ACM 48th International Conference on Software Engineering: Software Engineering in Practice. pp. 647–658. arXiv:2602.04449. doi:10.1145/3786583.3786904.
[edit]