Benchmarks for what frontier AI hasn't solved

Digital Work Index

An economically weighted measure of model performance across eight challenging benchmarks covering software engineering, logical reasoning and writing, and professional environments.
23.62
Fable 5.1
23.61
Opus 5.5
23.03
GPT-6 Astra
21.75
Opus 5
20.52
Grok 4.6
19.3
Grok 4.7
16.92
Gemini Flash 3.8
16.76
Kimi K3
16.45
GLM 5.3
16.29
Muse Spark 1.3
11.81
Qwen 3.8 Max
11.07
DeepSeek V4 Pro
6.95
Nemotron 3 Ultra 550B
All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Agentic Coding
Open Benchmarks Grants
Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
GPT-6 Astra
GPT-6 Astra
58.2%
2
Fable 5.1
Fable 5.1
57.9%
3
Opus 5
Opus 5
53.9%
Learn more about Terminal-Bench 4.0
Agentic Coding
Proprietary
Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
GPT-6 Astra
GPT-6 Astra
47.6%
2
Opus 5.5
Opus 5.5
40.4%
3
Fable 5.1
Fable 5.1
39.6%
Learn more about Agentic Coding 2.0
Agentic Coding
Open Benchmarks Grants
Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Opus 5
Opus 5
42.7%
2
GPT-5.6 Sol
GPT-5.6 Sol
34.6%
3
Fable 5
Fable 5
34.1%
Learn more about Terminal-Bench 3.0
View archived benchmarks
 Illution Back
Illution Front

Not more data. Better data.