Agentic Coding
Open Benchmarks Grants

Terminal-Bench 4.0

A benchmark to measure and evolve with the frontier of agent work: real terminal environments, real software engineering tasks, and a rolling task set that is revised as agents catch up to it. Version 4.0 trims the 3.0 task set from 74 to 66 tasks, removing eight and revising 20.

Built with
Image
Image
Image

At a glance

66

tasks in v4.0.0, down from 74 in v3.0.0

8

tasks removed in v4.0.0

20

tasks revised in v4.0.0

Terminal-Bench 4.0 Pareto Frontier

  • Opus 5.5
  • Sonnet 5.5
  • GPT-6 Astra
  • GPT-6.1 Sol
  • Fable 5.1
  • Opus 5
  • GPT-6 Sol
  • Fable 5
  • GLM-5.3
  • Grok 4.7
  • GPT-5.6 Sol
  • GLM-5.3-Flash
  • Qwen3.8-Max-0902
  • Opus 4.8
  • GPT-5.6 Terra
  • Grok 4.6
  • Gemini 3.8 Flash
  • GPT-5.6 Luna
  • GPT-6 Luna
  • Muse Spark 1.3
  • Grok 4.5
  • Sonnet 5
  • Gemini 3.7 Flash
Resolution Rate
Opus 5.5 ×
Sonnet 5.5 ×
GPT-6 Astra ×
GPT-6.1 Sol ×
Fable 5.1 ×
Opus 5 ×
GPT-6 Sol ×
Fable 5 ×
GLM-5.3 ×
Grok 4.7 ×
GPT-5.6 Sol ×
GLM-5.3-Flash ×
Qwen3.8-Max-0902 ×
Opus 4.8 ×
GPT-5.6 Terra ×
Grok 4.6 ×
Gemini 3.8 Flash ×
GPT-5.6 Luna ×
GPT-6 Luna ×
Muse Spark 1.3 ×
Grok 4.5 ×
Sonnet 5 ×
Gemini 3.7 Flash ×
Loading chart data...

Leaderboard

Rank Model Effort Agent Resolution Rate Release Date Tokens Cost
1 Opus 5.5 max Claude Code
64.8% ±3.1
2026-09-22 8.0B $4.7k
2 Sonnet 5.5 max Claude Code
61.8% ±2.9
2026-09-28 19.4B $7.3k
3 GPT-6 Astra max Codex
58.2% ±2.8
2026-09-03 1.5B $3.3k
4 GPT-6.1 Sol max Codex
58.2% ±3.1
2026-09-29 1.5B $0.6k
5 Fable 5.1 max Claude Code
57.9% ±3.8
2026-09-01 2.7B $6.2k
6 Opus 5 xhigh Claude Code
53.9% ±3.2
2026-07-24 6.9B $6.1k
7 GPT-6 Sol max Codex
49.4% ±3.2
2026-09-22 5.0B $1.5k
8 Fable 5 max Claude Code
44.5% ±3.8
2026-06-09 3.8B $7.3k
9 GLM-5.3 max Claude Code
41.8% ±3.2
2026-08-14 8.7B $2.7k
10 Grok 4.7 xhigh Grok Build
37.6% ±3.5
2026-09-21 5.5B $3.7k
11 GPT-5.6 Sol max Codex
37.3% ±3.8
2026-06-26 4.4B $2.5k
12 GLM-5.3-Flash none Claude Code
35.8% ±3.5
2026-08-26 2.9B $2.9k
13 Qwen3.8-Max-0902 max Claude Code
27% ±3.8
2026-09-02 3.0B $4.6k
14 Opus 4.8 max Claude Code
23.6% ±3.6
2026-05-28 6.4B $6.5k
15 GPT-5.6 Terra max Codex
21.5% ±3.3
2026-06-26 5.7B $1.7k
16 Grok 4.6 high Grok Build
20.3% ±3.1
2026-08-12 4.0B $3.6k
17 Gemini 3.8 Flash high mini-SWE-agent
19.1% ±3.4
2026-09-02 17.2B $1.8k
18 GPT-5.6 Luna max Codex
17.3% ±2.8
2026-06-26 11.6B $0.3k
19 GPT-6 Luna max Codex
16.4% ±2.7
2026-09-22 3.9B $0.1k
20 Muse Spark 1.3 xhigh Muse Code
14.5% ±2.9
2026-09-02 9.0B
21 Grok 4.5 high Grok Build
12.4% ±2.6
2026-07-16 3.4B $2.1k
22 Sonnet 5 max Claude Code
12.4% ±3.1
2026-06-30 21.6B $9.6k
23 Gemini 3.7 Flash high mini-SWE-agent
11.2% ±2.4
2026-08-13 11.1B $1.3k

How Terminal-Bench 4.0 compares

Benchmark

Released

Tasks

Change

Terminal-Bench 4.0

Aug 2026

66

8 tasks removed, 20 revised, 0 added

Terminal-Bench 3.0

Jul 2026

74

New rolling task set (formerly branded Frontier-Bench); initial 74 tasks added

Terminal-Bench 2.1

May 2026

89

Revision fixing 28 tasks; task count unchanged

Terminal-Bench 2.0

Nov 2025

89

Harder, curated task set; introduced Harbor, the agent eval and optimization framework

Terminal-Bench 1.0

May 2025

80

Initial release (Terminal-Bench-Core), led by Stanford and the Laude Institute

Methodology

Resolution rate

Share of the 66 v4.0.0 tasks each agent and model configuration resolves, reported with a margin reflecting run—to—run variance. “Effort” (max, high, and similar) is the harness’s own reasoning or agent effort setting for that run, not a Terminal-Bench parameter.
Cost and tokens
Totals are for the full evaluation run across all 66 tasks, not per task; they are not normalized for effort setting, so runs with a higher effort setting (like Sonnet 5’s) show substantially higher token and cost totals without a proportional gain in resolution rate.
Contamination control

The public Terminal-Bench site marks its task and leaderboard data with a canary string, requesting that benchmark data not appear in model training corpora.

Run command

Evaluated via the Harbor evaluation harness: harbor run -d terminal-bench/terminal-bench@4.0.0. Dataset pinned at hub.harborframework.com.

From the blog

Image for Terminal-Bench 4.0: Why Continuous Benchmarks Require Continuous QA

Terminal-Bench 4.0: Why Continuous Benchmarks Require Continuous QA

The speed of new frontier model releases keeps accelerating. Meanwhile benchmarks struggle to keep up and saturate quickly, often being...
August 28, 2026

Behind the benchmark

Terminal-Bench measures whether an AI agent can operate a real terminal to complete real software engineering and systems tasks: writing and debugging code, configuring environments, and recovering from failure, all without a GUI. Version 3.0 introduced the current 74 task families; version 4.0 is a maintenance release that removes eight tasks and revises twenty, netting 66 tasks, with no new tasks added in this cycle.

Acknowledgments

Terminal-Bench is hosted by Harbor and the Laude Institute, and supported by Snorkel AI via the Open Benchmarks Grants program. The v4.0.0 release was authored by Ryan Marten and Snorkel's Justin Bauer is among the benchmark's senior reviewers. 

FAQs

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

Enterprise Environments

MBABench 2.0

Evaluates agents on building financial models in spreadsheets, across 101 expert-written tasks.

By pass rate
1
Image
Astra (Codex)
15.8%
2
Image
Fable 5.1 (Claude Code)
7.9%
3
Image
Astra (ChatGPT)
7.9%
Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
GPT-6.1 Sol
42.8%
3
Image
Opus 5.5
40.4%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
GPT-6.1 Sol
19.6%
2
Image
Grok 4.6
16.4%
3
Image
Grok 4.7
12.1%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.