Frontier-Bench
The next frontier benchmark for terminal agents. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.
Overview
Terminal-Bench is used by virtually every frontier lab to measure whether agents can perform valuable work inside containerized terminal environments. Frontier-Bench is its official successor project designed to track the frontier with a diverse, difficult, high quality set of tasks that evolve over time, pushing past the ceiling of 2.1, where top agents already clear 75–84%.
Rather than a fixed release, Frontier-Bench is being assembled through open community contribution. Anyone can propose and submit a task; every submission passes through an automated review pipeline — static checks, a 35-criteria implementation rubric, Docker/oracle/no-op validation, live agent trials, and adversarial "cheat" trials — before a maintainer signs off.
Snorkel AI contributes to Frontier-Bench as both a task author and a data partner with additional support via the Open Benchmarks Grants program.
At a glance
74
tasks in the v0.1 release
7
31
subdomains spanning those 7 domains
35
rubric criteria every task must pass before merge
~34%
~34%
best model's pass rate on v0.1
Leaderboard
| Rank | Model | Agent | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Sol (max) | Codex |
34.4%
±1.6
|
2026-07-09 | 5.8B | $4.0k |
| 2 | Fable 5 (max) | Claude Code |
33.8%
±1.7
|
2026-06-09 | 3.6B | $7.0k |
| 3 | Opus 4.8 (max) | Claude Code |
21.1%
±1.6
|
2026-05-28 | 5.2B | $5.7k |
| 4 | GPT-5.6 Terra (max) | Codex |
20.8%
±1.4
|
2026-07-09 | 7.0B | $2.5k |
| 5 | Grok 4.5 (xhigh) | Cursor CLI |
17.8%
±1.4
|
2026-07-08 | 1.4B | $1.1k |
| 6 | Sonnet 5 (max) | Claude Code |
14.6%
±1.5
|
2026-06-30 | 17.9B | $7.9k |
| 7 | GPT-5.6 Luna (max) | Codex |
14.3%
±1.2
|
2026-07-09 | 12.0B | $1.7k |
| 8 | GLM 5.2 (max) | Claude Code |
5.1%
±1
|
2026-06-13 | 3.8B | $4.2k |
TAXONOMY
Seven domains, 31 subdomains
science
Natural sciences & engineering
Biology
Software
General software engineering
ML
Training, serving & eval
Operations
Business & financial reasoning
Security
Offensive & defensive security
Hardware
Physical & digital hardware
Creative & design work
Every task earns its place
Before a task counts toward Frontier-Bench, it passes through an automated + human review pipeline designed to catch broken tasks and reward-hacking before they ever reach the leaderboard.
How Frontier-Bench compares
Several recent benchmarks make progress on behavioral testing and instruction realism. The following table provides a brief comparison.
Benchmark
Task style
Domain span
Anti-cheat
Frontier-Bench
Community-submitted, expert-reviewed terminal tasks
7 domains / 31 subdomains
Rubric + cheat trials + hacker–fixer hardening loop
Terminal-Bench 2.1
89 fixed terminal tasks, continuously validated
Software, ML, security, data, science, sysadmin
Senior SWE-Bench
Real PRs from 12 OSS repos, taste-graded
Software engineering only
Rubric + bloat + practice + relative-taste gates
OSWorld 2.0
Long-horizon computer-use workflows
7 professional domains, 31 self-hosted sites
Separate safety audit (8 checks)
Resources
Acknowledgments
Frontier-Bench is hosted by Harbor and the Laude Institute, led by Ryan Marten, Alex Shaw, Andy Konwinski, and Ludwig Schmidt, with data and research contributions from Snorkel AI led by Justin Bauer and Vincent Sunn Chen, and additional support through Snorkel's Open Benchmarks Grants program.

