SWE-bench CLI
SWE-bench CLI evaluates whether an agent can complete substantive software engineering work inside an unfamiliar repository — inspecting the codebase, modifying multiple files, running tests, and leaving the repository in a verified state.
It spans ten technical domains, from backend systems and infrastructure to machine learning and security.
At a glance
200
frontier tasks
11
languages
10
technical domains
8
task types
Frontier performance
- Opus 5.5
- Fable 5.1
- Opus 5
- GPT-6 Astra
- Gemini Flash 3.8
- Grok 4.7
- Grok 4.6
- Qwen 3.8 Max
- Kimi K3
- GLM 5.3
- DeepSeek V4 Pro
- Muse Spark 1.3
- Nemotron 3 Ultra 550B
Key takeaways
SWE-bench CLI is a notably difficult benchmark, with Opus 5.5 leading at 16.6%, followed by Fable 5.1 at 14.5%, Opus 5 at 14.0%, and GPT-6 Astra at 12.4%. Every other model scored 6.3% or lower.
Failure patterns divided between unsuccessful verifier checks and interrupted execution. Agent timeouts affected more than 60% of runs for Grok 4.6 and GLM 5.3, and more than 85% for Muse Spark 1.3 and Qwen 3.8 Max.
Leaderboard
| Rank | Model | Pass@1 | Pass@5 | Cost / Trial | Cost / Task |
|---|---|---|---|---|---|
| 1 | Opus 5.5 |
16.6%
|
29.7%
|
$3.17 | $15.87 |
| 2 | Fable 5.1 |
14.5%
|
26.2%
|
$12.99 | $64.93 |
| 3 | Opus 5 |
14%
|
24.7%
|
$23.73 | $118.64 |
| 4 | GPT-6 Astra |
12.4%
|
18.8%
|
$2.16 | $10.8 |
| 5 | Gemini Flash 3.8 |
6.3%
|
14.1%
|
$4.61 | $23.06 |
| 6 | Grok 4.7 |
6.1%
|
17.1%
|
$5.2 | $25.99 |
| 7 | Grok 4.6 |
5.7%
|
13%
|
$3.3 | $16.5 |
| 8 | Qwen 3.8 Max |
4.9%
|
10.5%
|
$3.15 | $15.77 |
| 9 | Kimi K3 |
4.1%
|
11.9%
|
$6.09 | $30.45 |
| 10 | GLM 5.3 |
3.5%
|
10.8%
|
$8.19 | $40.96 |
| 11 | DeepSeek V4 Pro |
2.3%
|
5.2%
|
$2.2 | $10.98 |
| 12 | Muse Spark 1.3 |
2.2%
|
5.7%
|
$5.09 | $25.47 |
| 13 | Nemotron 3 Ultra 550B |
0.5%
|
1.1%
|
$4.62 | $23.11 |
Methodology
Evaluator
Each task includes a containerized repository, natural-language problem statement, reference patch, test configuration, and CLI-accessible development environment.
scoring note
The task reward is binary. A task receives credit only when all required tests pass. Golden changes must represent substantive engineering work spanning at least two files.
Behind the benchmark
Many coding evaluations reduce software engineering to generating a plausible patch. Real repository work requires navigating unfamiliar code, tracing behavior across files, selecting the right tests, diagnosing failures, and preserving existing functionality.
SWE-bench CLI focuses on that complete workflow. It measures whether an agent can operate as a software engineer inside a real codebase, not merely produce code that looks correct in isolation.

