LEADERBOARDS

Benchmarks for what frontier AI hasn't solved

Our ability to measure AI has been outpaced by our ability to develop it. We close that evaluation gap with benchmarks built around the tasks today's agents still break down on.
partners
Standford logo
Image
Image
Image
Image
Image
Image
Image
agent-le-logo
Image
Image
Cua Logo
Open Benchmarks Grants

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Built with
Image
Image
Image
Tasteful Solve Rate
The top-performing frontier models fail to complete tasks with senior-level correctness and taste over 65% of the time.
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
4
Image
GPT-5.6 Sol
34.7%
New
Open Benchmarks Grants

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
4
Image
GPT-6 Astra (Codex)
44.1%
Open Benchmarks Grants

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
4
Image
Opus 5
30.0%
Open Benchmarks Grants

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
4
Image
Fable 5
44.5%
Open Benchmarks Grants

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
4
Image
GLM 5.3
32.4%
Open Benchmarks Grants

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
4
Image
Opus 5 · medium
33.01%
Open Benchmarks Grants

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Claude Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
4
Image
Claude Opus 5
32.2%
Open Benchmarks Grants

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
4
Image
Opus 4.7
+0.183
Open Benchmarks Grants

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
4
Image
GPT 5.2
21.94%

View 9 archived benchmarks

Open Benchmarks Grants

Backed by a $3M commitment, our Open Benchmarks Grants program funds open-source datasets, benchmarks, and evaluation artifacts that shape how frontier AI is built and evaluated.

Get notified when we launch a new benchmark

Looking ahead

Three core dimensions where today's benchmarks fall short

Benchmarks must close the gap between what we measure and what agents actually encounter. Our work focuses on three dimensions where today’s evaluations break down.
01
Environment complexity
How dynamic is the operating environment? Real systems are far more complex than today's benchmarks.
02
Autonomy horizon
How independently can the agent operate before reliability breaks down?
03
Output complexity
How sophisticated is the deliverable agents must produce?
 Illution Back
Illution Front

For models that need to be right. Not just good enough.