Terminal-Bench 4.0
A benchmark to measure and evolve with the frontier of agent work: real terminal environments, real software engineering tasks, and a rolling task set that is revised as agents catch up to it. Version 4.0 trims the 3.0 task set from 74 to 66 tasks, removing eight and revising 20.
Overview
Terminal-Bench measures whether an AI agent can operate a real terminal to complete real software engineering and systems tasks: writing and debugging code, configuring environments, and recovering from failure, all without a GUI. Version 3.0 introduced the current 74 task families; version 4.0 is a maintenance release that removes eight tasks and revises twenty, netting 66 tasks, with no new tasks added in this cycle.
Headline finding, in this evaluation: Claude Opus 5 (Claude Code, max effort) leads at 51.8% ± 3.4% resolution, roughly 7 points ahead of Claude Fable 5 in second place. The bottom of the ranked field, Grok 4.5 and Claude Sonnet 5, both land at 12.4%, a spread of nearly 40 points across the ten evaluated model and agent configurations.
Snorkel AI contributes to Terminal-Bench 4.0 as both a task author and a data partner with additional support via the Open Benchmarks Grants program and Snorkel's Justin Bauer among the benchmark's reviewers.
At a glance
66
tasks in v4.0.0, down from 74 in v3.0.0
8
20
tasks revised in v4.0.0
Leaderboard
| Rank | Model | Effort | Agent | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|---|
| 1 | Opus 5 | max | Claude Code |
51.8%
±3.4
|
2026-07-24 | 6.5B | $6.0k |
| 2 | Fable 5 | max | Claude Code |
44.5%
±3.8
|
2026-06-09 | 3.8B | $7.3k |
| 3 | GLM-5.3 | max | Claude Code |
41.8%
±3.2
|
2026-08-14 | 8.7B | $2.7k |
| 4 | GPT-5.6 Sol | max | Codex |
37.3%
±3.8
|
2026-06-26 | 4.4B | $2.5k |
| 5 | Opus 4.8 | max | Claude Code |
23.6%
±3.6
|
2026-05-28 | 6.4B | $6.5k |
| 6 | GPT-5.6 Terra | max | Codex |
21.5%
±3.3
|
2026-06-26 | 5.7B | $1.7k |
| 7 | Grok 4.6 | high | Grok Build |
20.3%
±3.1
|
2026-08-12 | 4.0B | $3.6k |
| 8 | GPT-5.6 Luna | max | Codex |
17.3%
±2.8
|
2026-06-26 | 11.6B | $0.3k |
| 9 | Grok 4.5 | high | Grok Build |
12.4%
±2.6
|
2026-07-16 | 3.4B | $2.1k |
| 10 | Sonnet 5 | max | Claude Code |
12.4%
±3.1
|
2026-06-30 | 21.6B | $9.6k |
How Terminal-Bench 4.0 compares
Benchmark
Released
Tasks
Change
Terminal-Bench 4.0
Aug 2026
66
8 tasks removed, 20 revised, 0 added
Terminal-Bench 3.0
Jul 2026
74
New rolling task set (formerly branded Frontier-Bench); initial 74 tasks added
Terminal-Bench 2.1
May 2026
89
Revision fixing 28 tasks; task count unchanged
Terminal-Bench 2.0
Nov 2025
89
Harder, curated task set; introduced Harbor, the agent eval and optimization framework
Terminal-Bench 1.0
May 2025
80
Initial release (Terminal-Bench-Core), led by Stanford and the Laude Institute
Methodology
Resolution rate
The public Terminal-Bench site marks its task and leaderboard data with a canary string, requesting that benchmark data not appear in model training corpora.
Run command
Evaluated via the Harbor evaluation harness: harbor run -d terminal-bench/terminal-bench@4.0.0. Dataset pinned at hub.harborframework.com.
From the blog
Resources
Acknowledgments
Terminal-Bench is hosted by Harbor and the Laude Institute, and supported by Snorkel AI via the Open Benchmarks Grants program. The v4.0.0 release was authored by Ryan Marten and Snorkel's Justin Bauer is among the benchmark's senior reviewers.



