Terminal-Bench 3.0
The next frontier benchmark for terminal agents. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.
Overview
Terminal-Bench is used by virtually every frontier lab to measure whether agents can perform valuable work inside containerized terminal environments. Terminal-Bench 3.0 (formerly Frontier-Bench) is designed to track the frontier with a diverse, difficult, high quality set of tasks that evolve over time, pushing past the ceiling of 2.1, where top agents already clear 75–84%.
Rather than a fixed release, Terminal-Bench 3.0 is being assembled through open community contribution. Anyone can propose and submit a task; every submission passes through an automated review pipeline — static checks, a 35-criteria implementation rubric, Docker/oracle/no-op validation, live agent trials, and adversarial "cheat" trials — before a maintainer signs off.
Snorkel AI contributes to Terminal-Bench 3.0 as both a task author and a data partner with additional support via the Open Benchmarks Grants program and Snorkel's Justin Bauer among the benchmark's reviewers.
At a glance
74
tasks in the v0.1 release
7
31
subdomains spanning those 7 domains
35
rubric criteria every task must pass before merge
~43.5%
best model's pass rate on v0.1
Leaderboard
| Rank | Model | Agent | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 (max) | mini-SWE-agent |
43.5%
±1.7
|
2026-07-24 | 7.3B | $6.0k |
| 2 | GPT-5.6 Sol (max) | Codex |
34.4%
±1.6
|
2026-07-09 | 5.8B | $4.0k |
| 3 | Claude Fable 5 (max) | Claude Code |
33.8%
±1.7
|
2026-06-09 | 3.6B | $7.0k |
| 4 | Claude Opus 4.8 (max) | Claude Code |
21.1%
±1.6
|
2026-05-28 | 5.2B | $5.7k |
| 5 | GPT-5.6 Terra (max) | Codex |
20.8%
±1.4
|
2026-07-09 | 7.0B | $2.5k |
| 6 | Grok 4.5 (xhigh) | Cursor CLI |
17.8%
±1.4
|
2026-07-08 | 1.4B | $1.1k |
| 7 | Claude Sonnet 5 (max) | Claude Code |
14.6%
±1.5
|
2026-06-30 | 17.9B | $7.9k |
| 8 | GPT-5.6 Luna (max) | Codex |
14.3%
±1.2
|
2026-07-09 | 12.0B | $1.7k |
| 9 | GLM-5.2 (max) | Claude Code |
5.1%
±1
|
2026-06-13 | 3.8B | $4.2k |
TAXONOMY
Seven domains, 31 subdomains
science
Natural sciences & engineering
Biology
Software
General software engineering
ML
Training, serving & eval
Operations
Business & financial reasoning
Security
Offensive & defensive security
Hardware
Physical & digital hardware
Creative & design work
Sample tasks
Security
Shadow Relay
Reverse-engineer a network exfiltration from raw packet captures down to the decrypted payload, breaking a domain generation algorithm and a custom binary session along the way.
SOFTWARE
SOFTWARE
ML
Embedding Drift Monitor
SOFTWARE
Every task earns its place
Before a task counts toward Terminal-Bench 3.0, it passes through an automated + human review pipeline designed to catch broken tasks and reward-hacking before they ever reach the leaderboard.
How Terminal-Bench 3.0 compares
Several recent benchmarks make progress on behavioral testing and instruction realism. The following table provides a brief comparison.
Benchmark
Task style
Domain span
Anti-cheat
Terminal-Bench 3.0
Community-submitted, expert-reviewed terminal tasks
7 domains / 31 subdomains
Rubric + cheat trials + hacker–fixer hardening loop
Terminal-Bench 2.1
89 fixed terminal tasks, continuously validated
Software, ML, security, data, science, sysadmin
Senior SWE-Bench
Real PRs from 12 OSS repos, taste-graded
Software engineering only
Rubric + bloat + practice + relative-taste gates
OSWorld 2.0
Long-horizon computer-use workflows
7 professional domains, 31 self-hosted sites
Separate safety audit (8 checks)
Resources
Acknowledgments
Terminal-Bench 3.0 is hosted by Harbor and the Laude Institute, led by Ryan Marten, Alex Shaw, Andy Konwinski, and Ludwig Schmidt, with data and research contributions from Snorkel AI led by Justin Bauer and Vincent Sunn Chen, and additional support through Snorkel's Open Benchmarks Grants program. Justin Bauer was also among the benchmark's reviewers.

