Open Benchmarks Grants

Frontier-Bench

The next frontier benchmark for terminal agents. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.

Built with
Image
Image
Image

Overview

Terminal-Bench is used by virtually every frontier lab to measure whether agents can perform valuable work inside containerized terminal environments. Frontier-Bench is its official successor project designed to track the frontier with a diverse, difficult, high quality set of tasks that evolve over time, pushing past the ceiling of 2.1, where top agents already clear 75–84%.

Rather than a fixed release, Frontier-Bench is being assembled through open community contribution. Anyone can propose and submit a task; every submission passes through an automated review pipeline — static checks, a 35-criteria implementation rubric, Docker/oracle/no-op validation, live agent trials, and adversarial "cheat" trials — before a maintainer signs off.

Snorkel AI contributes to Frontier-Bench as both a task author and a data partner with additional support via the Open Benchmarks Grants program.

At a glance

74

tasks in the v0.1 release

7

top-level domains in the taxonomy

31

subdomains spanning those 7 domains

35

rubric criteria every task must pass before merge

~34%

best model's pass rate on v0.1

Leaderboard

Rank Model Agent Resolution Rate Release Date Tokens Cost
1 GPT-5.6 Sol (max) Codex
34.4% ±1.6
2026-07-09 5.8B $4.0k
2 Fable 5 (max) Claude Code
33.8% ±1.7
2026-06-09 3.6B $7.0k
3 Opus 4.8 (max) Claude Code
21.1% ±1.6
2026-05-28 5.2B $5.7k
4 GPT-5.6 Terra (max) Codex
20.8% ±1.4
2026-07-09 7.0B $2.5k
5 Grok 4.5 (xhigh) Cursor CLI
17.8% ±1.4
2026-07-08 1.4B $1.1k
6 Sonnet 5 (max) Claude Code
14.6% ±1.5
2026-06-30 17.9B $7.9k
7 GPT-5.6 Luna (max) Codex
14.3% ±1.2
2026-07-09 12.0B $1.7k
8 GLM 5.2 (max) Claude Code
5.1% ±1
2026-06-13 3.8B $4.2k

TAXONOMY

Seven domains, 31 subdomains

Every task is classified by the primary skill it exercises, not incidental tooling. The domain list is closed; subdomains grow as new tasks arrive.

science

Natural sciences & engineering

Biology

Chemistry
Physics
Earth
Robotics
Math
Linguistics

Software

General software engineering

Algorithms
Systems
Databases
Data engineering
Frontend
Languages

ML

Training, serving & eval

Training
Inference
Evaluation
Kernels

Operations

Business & financial reasoning

Finance
Logistics
Supply chain
Claims
Compliance
Marketing

Security

Offensive & defensive security

Cryptography
Reverse engineering
Forensics
AppSec

Hardware

Physical & digital hardware

CAD
RTL
Media

Creative & design work

Music
Design
Review pipeline

Every task earns its place

Before a task counts toward Frontier-Bench, it passes through an automated + human review pipeline designed to catch broken tasks and reward-hacking before they ever reach the leaderboard.

01

Static checks

Canary, Dockerfile sanity, path & metadata validation — no API keys required.
02
Rubric review
An agent scores the task against 35 criteria: verifiable, solvable, difficult, anti-cheat robust.
03
Docker / Oracle / No-op
Environment builds; the reference solution passes; doing nothing fails.
04

Maintainer review

A maintainer reviews the task once all automated gates pass.
05
Agent & cheat trials
Multiple agents attempt the task; adversarial trials probe for reward hacking.
06
Hacker–fixer loop
An adversarial loop iteratively hardens the task against exploits before merge.
07
Merged
Counts toward Frontier-Bench and becomes eligible for the leaderboard.
of

How Frontier-Bench compares

Several recent benchmarks make progress on behavioral testing and instruction realism. The following table provides a brief comparison.

Benchmark

Task style

Domain span

Anti-cheat

Frontier-Bench

Community-submitted, expert-reviewed terminal tasks

7 domains / 31 subdomains

Rubric + cheat trials + hacker–fixer hardening loop

Terminal-Bench 2.1

89 fixed terminal tasks, continuously validated

Software, ML, security, data, science, sysadmin

Continuous validation, community-reported fixes

Senior SWE-Bench

Real PRs from 12 OSS repos, taste-graded

Software engineering only

Rubric + bloat + practice + relative-taste gates

OSWorld 2.0

Long-horizon computer-use workflows

7 professional domains, 31 self-hosted sites

Separate safety audit (8 checks)

Acknowledgments

Frontier-Bench is hosted by Harbor and the Laude Institute, led by Ryan Marten, Alex Shaw, Andy Konwinski, and Ludwig Schmidt, with data and research contributions from Snorkel AI led by Justin Bauer and Vincent Sunn Chen, and additional support through Snorkel's Open Benchmarks Grants program.

FAQs

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

Senior SWE-bench

A benchmark for evaluating coding agents on senior-level engineering work: building features from realistic instructions, investigating bugs that require runtime investigation, and shipping code that aligns to existing codebase conventions.

Tasteful Solve Rate
1
Image
Claude Fable 5
29.1%
2
Image
Claude Opus 4.8
25.0%
3
Image
GPT-5.6 Sol
24.4%
Open Benchmarks Grants

OSWorld 2.0

A benchmark for evaluating computer-use agents on long-horizon, real-world workflows: 108 authentic tasks across 31 self-hosted web environments and professional desktop applications.

By Binary Accuracy (300 steps)
1
Image
gpt-5-5 · xhigh
13%
2
Image
Claude Opus 4.7 · max
13%
3
Image
Claude Sonnet 4.6 · medium
8.3%
Open Benchmarks Grants

Agents’ Last Exam

Long-horizon professional workflows with verifiable outcomes across 55 sub-industries. 147 public tasks of a 1,500+ task corpus, sourced and validated by 300+ industry experts.

Top Submissions (Pass Rate)
1
Image
Codex · GPT 5.6-Sol · XHigh
30.6%
2
Image
Codex · GPT 5.6-Sol · High
30.6%
3
Image
Codex · GPT 5.6-Sol · Max
29.6%

Agentic Coding

A benchmark for evaluating AI models on complex, real-world coding tasks that require multi-step reasoning, tool use, and autonomous problem-solving.

Top Models
1
Image
Claude Opus 4.6
65.2%
2
Image
Claude Opus 4.5
58.0%
3
Image
Claude Sonnet 4.5
57.6%
Open Benchmarks Grants

SlopCode Bench

Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.

Top Models by Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3-Codex
26.02%
3
Image
GPT-5.4
23.47%
Open Benchmarks Grants

Continual Learning Bench

Evaluates whether AI systems improve from prior experience across sequential, stateful tasks, measuring real in-context learning, not just raw capability.

Top Systems (Agg. Reward)
1
Image
ICL · Claude Sonnet 4.6
+0.196
2
Image
ICL · GPT-5.4
+0.189
3
Image
Claude Code · Sonnet 4.6
+0.185
Open Benchmarks Grants

Terminal-Bench 2.1

Terminal agent evaluation led by Stanford University and Laude Institute. v2.1 fixes 28 tasks from 2.0 and introduces continuous validation.

Top Submissions
1
Image
Codex CLI · GPT-5.5
83.4%
2
Image
Claude Code · Claude 5 Fable
83.1%
3
Image
Terminus 2 · Claude 5 Fable
80.4%
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.