Software Engineering

SWE-bench CLI

SWE-bench CLI evaluates whether an agent can complete substantive software engineering work inside an unfamiliar repository — inspecting the codebase, modifying multiple files, running tests, and leaving the repository in a verified state.

It spans ten technical domains, from backend systems and infrastructure to machine learning and security.

At a glance

200

frontier tasks

11

languages

10

technical domains

8

task types

Frontier performance

  • Opus 5.5
  • Fable 5.1
  • Opus 5
  • GPT-6 Astra
  • Gemini Flash 3.8
  • Grok 4.7
  • Grok 4.6
  • Qwen 3.8 Max
  • Kimi K3
  • GLM 5.3
  • DeepSeek V4 Pro
  • Muse Spark 1.3
  • Nemotron 3 Ultra 550B
Opus 5.5 ×
Fable 5.1 ×
Opus 5 ×
GPT-6 Astra ×
Gemini Flash 3.8 ×
Grok 4.7 ×
Grok 4.6 ×
Qwen 3.8 Max ×
Kimi K3 ×
GLM 5.3 ×
DeepSeek V4 Pro ×
Muse Spark 1.3 ×
Nemotron 3 Ultra 550B ×
Loading chart data...

Key takeaways

SWE-bench CLI is a notably difficult benchmark, with Opus 5.5 leading at 16.6%, followed by Fable 5.1 at 14.5%, Opus 5 at 14.0%, and GPT-6 Astra at 12.4%. Every other model scored 6.3% or lower.

Failure patterns divided between unsuccessful verifier checks and interrupted execution. Agent timeouts affected more than 60% of runs for Grok 4.6 and GLM 5.3, and more than 85% for Muse Spark 1.3 and Qwen 3.8 Max.

Leaderboard

Rank Model Pass@1 Pass@5 Cost / Trial Cost / Task
1 Opus 5.5
16.6%
29.7%
$3.17 $15.87
2 Fable 5.1
14.5%
26.2%
$12.99 $64.93
3 Opus 5
14%
24.7%
$23.73 $118.64
4 GPT-6 Astra
12.4%
18.8%
$2.16 $10.8
5 Gemini Flash 3.8
6.3%
14.1%
$4.61 $23.06
6 Grok 4.7
6.1%
17.1%
$5.2 $25.99
7 Grok 4.6
5.7%
13%
$3.3 $16.5
8 Qwen 3.8 Max
4.9%
10.5%
$3.15 $15.77
9 Kimi K3
4.1%
11.9%
$6.09 $30.45
10 GLM 5.3
3.5%
10.8%
$8.19 $40.96
11 DeepSeek V4 Pro
2.3%
5.2%
$2.2 $10.98
12 Muse Spark 1.3
2.2%
5.7%
$5.09 $25.47
13 Nemotron 3 Ultra 550B
0.5%
1.1%
$4.62 $23.11

Methodology

Evaluator

Harbor runs repository-level fail-to-pass and pass-to-pass test suites against the agent’s resulting code. Task metadata and rubric checks provide additional requirements for correctness.
timeout
Task-specific agent and verifier timeouts bound each CLI trace. All evaluations run in reproducible, network-disabled Docker environments.
integration

Each task includes a containerized repository, natural-language problem statement, reference patch, test configuration, and CLI-accessible development environment.

scoring note

The task reward is binary. A task receives credit only when all required tests pass. Golden changes must represent substantive engineering work spanning at least two files.

Behind the benchmark

Many coding evaluations reduce software engineering to generating a plausible patch. Real repository work requires navigating unfamiliar code, tracing behavior across files, selecting the right tests, diagnosing failures, and preserving existing functionality.

SWE-bench CLI focuses on that complete workflow. It measures whether an agent can operate as a software engineer inside a real codebase, not merely produce code that looks correct in isolation.

More benchmarks

Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Opus 5.5
40.4%
3
Image
Fable 5.1
39.6%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
Grok 4.7
12.1%
3
Image
GLM 5.3
11.9%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.