Agentic Coding
Open Benchmarks Grants

Terminal-Bench 3.0

The next frontier benchmark for terminal agents. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.

Built with
Image
Image
Image
Image

At a glance

74

tasks in the v0.1 release

7

top-level domains in the taxonomy

31

subdomains spanning those 7 domains

35

rubric criteria every task must pass before merge

~34%

best model's pass rate
on v0.1, vs. ~5% for
the best open-weight model

Leaderboard

Rank Model Effort Agent Resolution Rate Release Date Tokens Cost
1 Opus 5 max mini-SWE-agent
42.7% ±1.6
2026-07-24 7.3B $5.8k
2 GPT-5.6 Sol max Codex
34.6% ±1.6
2026-07-09 5.8B $4.0k
3 Fable 5 max Claude Code
34.1% ±1.7
2026-06-09 3.6B $6.5k
4 GLM 5.3 max Claude Code
32.4% ±1.5
2026-08-14 5.6B $1.8k
5 Grok 4.6 high Grok Build
26.5% ±1.5
2026-08-12 2.9B $2.1k
6 Opus 4.8 max Claude Code
21.1% ±1.6
2026-05-28 5.2B $5.2k
7 GPT-5.6 Terra max Codex
20.8% ±1.4
2026-07-09 7.0B $2.5k
8 SWE-1.7 Lightning Devin
18.6% ±1.5
2026-07-08 3.6B $7.2k
9 Grok 4.5 xhigh Cursor CLI
15.7% ±1.5
2026-07-08 1.2B $766.02
10 Sonnet 5 max Claude Code
14.6% ±1.5
2026-06-30 17.9B $6.9k
11 GPT-5.6 Luna max Codex
14.3% ±1.3
2026-07-09 11.9B $1.6k
12 GLM 5.2 max Claude Code
4.6% ±1
2026-06-13 3.3B $3.4k

TAXONOMY

Seven domains, 31 subdomains

Every task is classified by the primary skill it exercises, not incidental tooling. The domain list is closed; subdomains grow as new tasks arrive.

science

Natural sciences & engineering

Biology

Chemistry
Physics
Earth
Robotics
Math
Linguistics

Software

General software engineering

Algorithms
Systems
Databases
Data engineering
Frontend
Languages

ML

Training, serving & eval

Training
Inference
Evaluation
Kernels

Operations

Business & financial reasoning

Finance
Logistics
Supply chain
Claims
Compliance
Marketing

Security

Offensive & defensive security

Cryptography
Reverse engineering
Forensics
AppSec

Hardware

Physical & digital hardware

CAD
RTL
Media

Creative & design work

Music
Design
Review pipeline

Every task earns its place

Before a task counts toward Terminal-Bench 3.0, it passes through an automated + human review pipeline designed to catch broken tasks and reward-hacking before they ever reach the leaderboard.

01

Static checks

Canary, Dockerfile sanity, path & metadata validation — no API keys required.
02
Rubric review
An agent scores the task against 35 criteria: verifiable, solvable, difficult, anti-cheat robust.
03
Docker / Oracle / No-op
Environment builds; the reference solution passes; doing nothing fails.
04

Maintainer review

A maintainer reviews the task once all automated gates pass.
05
Agent & cheat trials
Multiple agents attempt the task; adversarial trials probe for reward hacking.
06
Hacker–fixer loop
An adversarial loop iteratively hardens the task against exploits before merge.
07
Merged
Counts toward Frontier-Bench and becomes eligible for the leaderboard.
of

How Terminal-Bench 3.0 compares

Several recent benchmarks make progress on behavioral testing and instruction realism. The following table provides a brief comparison.

Benchmark

Task style

Domain span

Anti-cheat

Terminal-Bench 3.0

Community-submitted, expert-reviewed terminal tasks

7 domains / 31 subdomains

Rubric + cheat trials + hacker–fixer hardening loop

Terminal-Bench 2.1

89 fixed terminal tasks, continuously validated

Software, ML, security, data, science, sysadmin

Continuous validation, community-reported fixes

Senior SWE-Bench

Real PRs from 12 OSS repos, taste-graded

Software engineering only

Rubric + bloat + practice + relative-taste gates

OSWorld 2.0

Long-horizon computer-use workflows

7 professional domains, 31 self-hosted sites

Separate safety audit (8 checks)

From the blog

Image for Why Frontier Agents Fail Real Engineering Work: Two Terminal-Bench 3.0 Task Deep Dives

Why Frontier Agents Fail Real Engineering Work: Two Terminal-Bench 3.0 Task Deep Dives

Terminal-Bench 3.0 (formerly Frontier-Bench) recently launched, built to track what AI agents can and can’t do across real computer work....

Behind the benchmark

Terminal-Bench 3.0 is hosted by Harbor and the Laude Institute, led by Ryan Marten, Alex Shaw, Andy Konwinski, and Ludwig Schmidt, with data and research contributions from Snorkel AI led by Justin Bauer and Vincent Sunn Chen, and additional support through Snorkel's Open Benchmarks Grants program. Justin Bauer was also among the benchmark's reviewers.

Acknowledgments

Terminal-Bench 3.0 is hosted by Harbor and the Laude Institute, led by Ryan Marten, Alex Shaw, Andy Konwinski, and Ludwig Schmidt, with data and research contributions from Snorkel AI led by Justin Bauer and Vincent Sunn Chen, and additional support through Snorkel's Open Benchmarks Grants program. Justin Bauer was also among the benchmark's reviewers.

FAQs

Get notified when we launch a new benchmark

Share this benchmark
 Illution Back
Illution Front

For models that need to be right. Not just good enough.