Agentic Coding 2.0
A frontier benchmark for evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.
Built from a representative frontier subset of our Terminal-Bench+ dataset, it expands the evaluation from difficult coding problems to autonomous engineering.
At a glance
200
frontier tasks
9
tasks types
9
target languages
Frontier performance
- GPT-6 Astra
- Opus 5.5
- Fable 5.1
- Opus 5
- Grok 4.6
- Gemini Flash 3.8
- Grok 4.7
- Kimi K3
- Muse Spark 1.3
- GLM 5.3
- DeepSeek V4 Pro
- Qwen 3.8 Max
- Nemotron 3 Ultra 550B
Key takeaways
GPT-6 Astra led Agentic Coding 2.0 at 47.6% Pass@1, followed by Opus 5.5 at 40.4%, Fable 5.1 at 39.6%, and Opus 5 at 38.9%. Astra's largest advantages appeared in build and dependency management at 67.5% and security at 53.8%.
Execution reliability created substantial separation across the field. Agent-timeout rates ranged from zero for Astra to 77.7% for Qwen 3.8 Max.
Leaderboard
| Rank | Model | Pass@1 | Pass@5 | Cost / Trial | Cost / Task |
|---|---|---|---|---|---|
| 1 | GPT-6 Astra |
47.6%
|
58.7%
|
$1.16 | $5.79 |
| 2 | Opus 5.5 |
40.4%
|
56.9%
|
$0.94 | $4.7 |
| 3 | Fable 5.1 |
39.6%
|
54.8%
|
$3.86 | $19.28 |
| 4 | Opus 5 |
38.9%
|
62%
|
$2.91 | $14.55 |
| 5 | Grok 4.6 |
33.6%
|
53.1%
|
$1.11 | $5.57 |
| 6 | Gemini Flash 3.8 |
32.6%
|
53.1%
|
$1.56 | $7.8 |
| 7 | Grok 4.7 |
31.9%
|
56%
|
$1.61 | $8.03 |
| 8 | Kimi K3 |
26.1%
|
52.4%
|
$1.58 | $7.9 |
| 9 | Muse Spark 1.3 |
25%
|
40.3%
|
$1.83 | $9.13 |
| 10 | GLM 5.3 |
24.9%
|
46.3%
|
$1.2 | $6.01 |
| 11 | DeepSeek V4 Pro |
14.8%
|
35.6%
|
$1.66 | $8.29 |
| 12 | Qwen 3.8 Max |
14.7%
|
26.2%
|
$0.9 | $4.49 |
| 13 | Nemotron 3 Ultra 550B |
5.1%
|
10.6%
|
$0.92 | $4.6 |
Methodology
Evaluator
Tasks use the Harbor Terminal-Bench format with Docker environments, packaged dependencies, reference solutions, supporting files, and no network access.
scoring note
Five attempts per task support pass-rate calibration and repeatability. A task must satisfy its deterministic tests and required rubric criteria to receive credit.
Behind the benchmark
Agentic Coding 2.0 is the next generation of Snorkel’s original Agentic Coding benchmark. Built from a representative frontier subset of our Terminal-Bench+ dataset, it expands the evaluation from difficult coding problems to autonomous engineering across software engineering, debugging, systems, security, data processing, machine learning, scientific computing, games, and build and dependency management.
Each task runs in an isolated, air-gapped environment. Agents must use tools, manage intermediate state, validate their work, and recover from errors. Every task includes a reference solution, deterministic tests, and rubrics that evaluate both the final result and the agent’s trajectory.
The original Agentic Coding benchmark established a focused evaluation of multi-step coding tasks. Agentic Coding 2.0 broadens that foundation into a more demanding test of autonomous engineering.
The benchmark evaluates the complete agent loop: understanding the assignment, exploring an unfamiliar environment, selecting and sequencing tools, executing a solution, inspecting the result, correcting mistakes, and producing a verifiable final state.

