Agentic Coding

Agentic Coding 2.0

A frontier benchmark for evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

Built from a representative frontier subset of our Terminal-Bench+ dataset, it expands the evaluation from difficult coding problems to autonomous engineering.

At a glance

200

frontier tasks

9

tasks types

9

target languages

Frontier performance

  • GPT-6 Astra
  • Opus 5.5
  • Fable 5.1
  • Opus 5
  • Grok 4.6
  • Gemini Flash 3.8
  • Grok 4.7
  • Kimi K3
  • Muse Spark 1.3
  • GLM 5.3
  • DeepSeek V4 Pro
  • Qwen 3.8 Max
  • Nemotron 3 Ultra 550B
GPT-6 Astra ×
Opus 5.5 ×
Fable 5.1 ×
Opus 5 ×
Grok 4.6 ×
Gemini Flash 3.8 ×
Grok 4.7 ×
Kimi K3 ×
Muse Spark 1.3 ×
GLM 5.3 ×
DeepSeek V4 Pro ×
Qwen 3.8 Max ×
Nemotron 3 Ultra 550B ×
Loading chart data...

Key takeaways

GPT-6 Astra led Agentic Coding 2.0 at 47.6% Pass@1, followed by Opus 5.5 at 40.4%, Fable 5.1 at 39.6%, and Opus 5 at 38.9%. Astra's largest advantages appeared in build and dependency management at 67.5% and security at 53.8%.

Execution reliability created substantial separation across the field. Agent-timeout rates ranged from zero for Astra to 77.7% for Qwen 3.8 Max.

Leaderboard

Rank Model Pass@1 Pass@5 Cost / Trial Cost / Task
1 GPT-6 Astra
47.6%
58.7%
$1.16 $5.79
2 Opus 5.5
40.4%
56.9%
$0.94 $4.7
3 Fable 5.1
39.6%
54.8%
$3.86 $19.28
4 Opus 5
38.9%
62%
$2.91 $14.55
5 Grok 4.6
33.6%
53.1%
$1.11 $5.57
6 Gemini Flash 3.8
32.6%
53.1%
$1.56 $7.8
7 Grok 4.7
31.9%
56%
$1.61 $8.03
8 Kimi K3
26.1%
52.4%
$1.58 $7.9
9 Muse Spark 1.3
25%
40.3%
$1.83 $9.13
10 GLM 5.3
24.9%
46.3%
$1.2 $6.01
11 DeepSeek V4 Pro
14.8%
35.6%
$1.66 $8.29
12 Qwen 3.8 Max
14.7%
26.2%
$0.9 $4.49
13 Nemotron 3 Ultra 550B
5.1%
10.6%
$0.92 $4.6

Methodology

Evaluator

Harbor evaluates deterministic task tests, required outputs, intermediate milestones, and trace-level behavior. Frontier difficulty is calibrated through repeated runs of frontier coding agents.
timeout
Each task has bounded agent, verifier, and environment-build limits. Agent and verifier traces are limited to 30 minutes.
integration

Tasks use the Harbor Terminal-Bench format with Docker environments, packaged dependencies, reference solutions, supporting files, and no network access.

scoring note

Five attempts per task support pass-rate calibration and repeatability. A task must satisfy its deterministic tests and required rubric criteria to receive credit.

Behind the benchmark

Agentic Coding 2.0 is the next generation of Snorkel’s original Agentic Coding benchmark. Built from a representative frontier subset of our Terminal-Bench+ dataset, it expands the evaluation from difficult coding problems to autonomous engineering across software engineering, debugging, systems, security, data processing, machine learning, scientific computing, games, and build and dependency management.

Each task runs in an isolated, air-gapped environment. Agents must use tools, manage intermediate state, validate their work, and recover from errors. Every task includes a reference solution, deterministic tests, and rubrics that evaluate both the final result and the agent’s trajectory.

The original Agentic Coding benchmark established a focused evaluation of multi-step coding tasks. Agentic Coding 2.0 broadens that foundation into a more demanding test of autonomous engineering.

The benchmark evaluates the complete agent loop: understanding the assignment, exploring an unfamiliar environment, selecting and sequencing tools, executing a solution, inspecting the result, correcting mistakes, and producing a verifiable final state.

More benchmarks

Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
Grok 4.7
12.1%
3
Image
GLM 5.3
11.9%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.