Terminal-Bench 3.0
The next frontier benchmark for terminal agents. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.
At a glance
74
7
31
35
~34%
best model's pass rate
on v0.1, vs. ~5% for
the best open-weight model
Leaderboard
| Rank | Model | Effort | Agent | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|---|
| 1 | Opus 5 | max | mini-SWE-agent |
42.7%
±1.6
|
2026-07-24 | 7.3B | $5.8k |
| 2 | GPT-5.6 Sol | max | Codex |
34.6%
±1.6
|
2026-07-09 | 5.8B | $4.0k |
| 3 | Fable 5 | max | Claude Code |
34.1%
±1.7
|
2026-06-09 | 3.6B | $6.5k |
| 4 | GLM 5.3 | max | Claude Code |
32.4%
±1.5
|
2026-08-14 | 5.6B | $1.8k |
| 5 | Grok 4.6 | high | Grok Build |
26.5%
±1.5
|
2026-08-12 | 2.9B | $2.1k |
| 6 | Opus 4.8 | max | Claude Code |
21.1%
±1.6
|
2026-05-28 | 5.2B | $5.2k |
| 7 | GPT-5.6 Terra | max | Codex |
20.8%
±1.4
|
2026-07-09 | 7.0B | $2.5k |
| 8 | SWE-1.7 Lightning | Devin |
18.6%
±1.5
|
2026-07-08 | 3.6B | $7.2k | |
| 9 | Grok 4.5 | xhigh | Cursor CLI |
15.7%
±1.5
|
2026-07-08 | 1.2B | $766.02 |
| 10 | Sonnet 5 | max | Claude Code |
14.6%
±1.5
|
2026-06-30 | 17.9B | $6.9k |
| 11 | GPT-5.6 Luna | max | Codex |
14.3%
±1.3
|
2026-07-09 | 11.9B | $1.6k |
| 12 | GLM 5.2 | max | Claude Code |
4.6%
±1
|
2026-06-13 | 3.3B | $3.4k |
TAXONOMY
Seven domains, 31 subdomains
science
Natural sciences & engineering
Biology
Software
General software engineering
ML
Training, serving & eval
Operations
Business & financial reasoning
Security
Offensive & defensive security
Hardware
Physical & digital hardware
Creative & design work
Sample tasks
Security
Shadow Relay
Reverse-engineer a network exfiltration from raw packet captures down to the decrypted payload, breaking a domain generation algorithm and a custom binary session along the way.
SOFTWARE
SOFTWARE
ML
Embedding Drift Monitor
SOFTWARE
Every task earns its place
Before a task counts toward Terminal-Bench 3.0, it passes through an automated + human review pipeline designed to catch broken tasks and reward-hacking before they ever reach the leaderboard.
How Terminal-Bench 3.0 compares
Several recent benchmarks make progress on behavioral testing and instruction realism. The following table provides a brief comparison.
Benchmark
Task style
Domain span
Anti-cheat
Terminal-Bench 3.0
Community-submitted, expert-reviewed terminal tasks
7 domains / 31 subdomains
Rubric + cheat trials + hacker–fixer hardening loop
Terminal-Bench 2.1
89 fixed terminal tasks, continuously validated
Software, ML, security, data, science, sysadmin
Senior SWE-Bench
Real PRs from 12 OSS repos, taste-graded
Software engineering only
Rubric + bloat + practice + relative-taste gates
OSWorld 2.0
Long-horizon computer-use workflows
7 professional domains, 31 self-hosted sites
Separate safety audit (8 checks)
From the blog
Resources
Behind the benchmark
Behind the benchmark
Terminal-Bench 3.0 is hosted by Harbor and the Laude Institute, led by Ryan Marten, Alex Shaw, Andy Konwinski, and Ludwig Schmidt, with data and research contributions from Snorkel AI led by Justin Bauer and Vincent Sunn Chen, and additional support through Snorkel's Open Benchmarks Grants program. Justin Bauer was also among the benchmark's reviewers.
Acknowledgments
Terminal-Bench 3.0 is hosted by Harbor and the Laude Institute, led by Ryan Marten, Alex Shaw, Andy Konwinski, and Ludwig Schmidt, with data and research contributions from Snorkel AI led by Justin Bauer and Vincent Sunn Chen, and additional support through Snorkel's Open Benchmarks Grants program. Justin Bauer was also among the benchmark's reviewers.




