LEADERBOARDS

Benchmarks for what frontier AI hasn't solved

Our ability to measure AI has been outpaced by our ability to develop it. We close that evaluation gap with benchmarks built around the tasks today's agents still break down on.
partners
Standford logo
Image
Image
Image
Image
Image
Image
Image
agent-le-logo
Image
Image
Cua Logo

Senior SWE-bench

A benchmark for evaluating coding agents on senior-level engineering work: building features from realistic instructions, investigating bugs that require runtime investigation, and shipping code that aligns to existing codebase conventions.

Built with
Image
Image
Image
Tasteful Solve Rate
The top-performing frontier models fail to complete tasks with senior-level correctness and taste over 65% of the time.
1
Image
Claude Fable 5.1
34.7%
2
Image
Claude Fable 5
34.7%
3
Image
Claude Opus 5
34.7%
4
Image
GPT-5.6 Sol
34.7%
New
Open Benchmarks Grants

Terminal-Bench 4.0

A benchmark to measure and evolve with the frontier of agent work: real terminal environments, real software engineering tasks, and a rolling task set that is revised as agents catch up to it.

By Resolution Rate
1
Image
GPT-6 Astra (max) • Codex
58.2%
2
Image
Fable 5.1 (max) • Claude Code
57.9%
3
Image
Opus 5 (max) • Claude Code
51.8%
Open Benchmarks Grants

Terminal-Bench-Science

A benchmark for evaluating AI agents on workflows from researchers’ own work. Scientists, not model developers or vendors, set the bar for scientific capability in AI.

By Resolution Rate
1
Image
Opus 5 (max) • Claude Code
30.0%
2
Image
GPT-5.6 Sol (max) • Codex
22.4%
3
Image
Fable 5 (max) • Claude Code
21.4%
Open Benchmarks Grants

Terminal-Bench 3.0

The next frontier benchmark for agent work. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.

By Resolution Rate
1
Image
Opus 5 (max) • mini-SWE-agent
42.7%
2
Image
GPT-5.6 Sol (max) • Codex
34.6%
3
Image
Fable 5 (max) • Claude Code
34.1%
Open Benchmarks Grants

OSWorld 2.0

A benchmark for evaluating computer-use agents on long-horizon, real-world workflows: 108 authentic tasks across 31 self-hosted web environments and professional desktop applications.

By Binary Accuracy (500 steps)
1
Image
Claude Opus 5 (max, Batch tool)
31.43%
2
Image
GPT-5.6 Sol (max, Batch tool)
27.34%
3
Image
Claude Opus 4.8 (max, Batched tool)
20.6%
Open Benchmarks Grants

Agents’ Last Exam

Long-horizon professional workflows with verifiable outcomes across 55 sub-industries. 147 public tasks of a 1,500+ task corpus, sourced and validated by 300+ industry experts.

Top Submissions (Pass Rate)
1
Image
Claude Code · Claude Opus 5 · High
31.6%
2
Image
Codex · GPT 5.6-Sol · XHigh
30.6%
3
Image
Codex · GPT 5.6-Luna · XHigh
30.3%
Open Benchmarks Grants

Continual Learning Bench

Evaluates whether AI systems improve from prior experience across sequential, stateful tasks, measuring real in-context learning, not just raw capability.

Top Systems (Agg. Reward)
1
Image
ICL · Claude Sonnet 4.6
+0.196
2
Image
ICL · GPT-5.4
+0.189
3
Image
Claude Code · Sonnet 4.6
+0.185
Open Benchmarks Grants

SlopCode Bench

Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.

Top Models by Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3-Codex
26.02%
3
Image
GPT-5.4
23.47%

View 9 archived benchmarks

Open Benchmarks Grants

Backed by a $3M commitment, our Open Benchmarks Grants program funds open-source datasets, benchmarks, and evaluation artifacts that shape how frontier AI is built and evaluated.

Get notified when we launch a new benchmark

Looking ahead

Three core dimensions where today's benchmarks fall short

Benchmarks must close the gap between what we measure and what agents actually encounter. Our work focuses on three dimensions where today’s evaluations break down.
01
Environment complexity
How dynamic is the operating environment? Real systems are far more complex than today's benchmarks.
02
Autonomy horizon
How independently can the agent operate before reliability breaks down?
03
Output complexity
How sophisticated is the deliverable agents must produce?
 Illution Back
Illution Front

For models that need to be right. Not just good enough.