Datasets & Benchmarks · 2023
SWE-bench
Princeton asked whether AI could fix real GitHub bugs; in 2023 the best model managed under 2 percent.
By 2023 language models were acing toy coding puzzles, so Princeton researchers Carlos Jimenez, John Yang and colleagues built a benchmark from the job itself. SWE-bench collected 2,294 real issue and pull request pairs from twelve popular open source Python projects, including Django, scikit-learn and matplotlib.
The task mirrors a working engineer's day: given a full repository and the text of a bug report or feature request, produce a patch. Success is judged by the project's own test suite, using tests that fail before the fix and pass after it, so partial credit and plausible-looking answers count for nothing.
The first results were sobering. Claude 2, the strongest model tested, resolved just 1.96 percent of issues, and others did worse. Real codebases demanded skills the puzzle benchmarks never touched: navigating thousands of files, understanding cross-module context and editing without breaking distant behavior.
That difficulty made SWE-bench the engine of the coding agent era. Systems learned to read stack traces, search repositories, run tests and iterate, and scaffolds like SWE-agent in 2024 showed that tooling around the model mattered as much as the model. OpenAI later released SWE-bench Verified, a human-screened subset of 500 tasks that fixed broken or unfair cases and became the standard variant.
Scores rose from under 2 percent to above 70 percent on Verified within about two years, one of the steepest capability climbs ever recorded on a benchmark. Debates continue about contamination and its Python-only scope, but SWE-bench remains the number enterprises watch to judge whether coding agents are ready for production work.
Related stories
From history to production
We turn these ideas into working systems
The same techniques, shipped into your stack with evals, observability, and measurable ROI.