Benchmarks for what frontier AI hasn't solved
Digital Work Index
An economically weighted measure of model performance across eight challenging benchmarks covering software engineering, logical reasoning and writing, and professional environments.
23.62
Fable 5.1
23.61
Opus 5.5
23.03
GPT-6 Astra
21.75
Opus 5
20.52
Grok 4.6
19.3
Grok 4.7
16.92
Gemini Flash 3.8
16.76
Kimi K3
16.45
GLM 5.3
16.29
Muse Spark 1.3
11.81
Qwen 3.8 Max
11.07
DeepSeek V4 Pro
6.95
Nemotron 3 Ultra 550B
All Models
Sep 22, 2026
Opus 5.5
RANK BY BENCHMARK
Agentic Coding 2.0
40.4%
#2/13
Agents’ Last Exam
38.2%
#1/43
SnorkelFinance 2.0
18.6%
#5/13
SnorkelLegal
31.4%
#7/13
SnorkelManufacturing
10.4%
#6/13
Sep 21, 2026
Grok 4.7
RANK BY BENCHMARK
Agentic Coding 2.0
31.9%
#7/13
Senior SWE-Bench
27.4%
#8/19
SnorkelFinance 2.0
20.2%
#3/13
SnorkelLegal
32.8%
#6/13
SnorkelManufacturing
12.1%
#2/13
Sep 03, 2026
GPT-6 Astra
RANK BY BENCHMARK
Agentic Coding 2.0
47.6%
#1/13
Agents’ Last Exam
34.2%
#2/43
SnorkelFinance 2.0
11.6%
#9/13
SnorkelLegal
20.9%
#13/13
SnorkelManufacturing
10%
#8/13
Sep 02, 2026
Gemini Flash 3.8
RANK BY BENCHMARK
Agentic Coding 2.0
32.6%
#6/13
SnorkelFinance 2.0
5.4%
#12/13
SnorkelLegal
25.3%
#12/13
SnorkelManufacturing
7.1%
#12/13
SnorkelRevOps
5.4%
#13/13
All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Agentic Coding
Open Benchmarks Grants
Terminal-Bench 4.0
Evaluates terminal agents on continuously updated software-engineering tasks.
By Resolution Rate
1
GPT-6 Astra
58.2%
2
Fable 5.1
57.9%
3
Opus 5
53.9%
Agentic Coding
Proprietary
Agentic Coding 2.0
Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.
By pass@1
1
GPT-6 Astra
47.6%
2
Opus 5.5
40.4%
3
Fable 5.1
39.6%
Agentic Coding
Open Benchmarks Grants
Terminal-Bench 3.0
Evaluates terminal agents on containerized tasks across seven domains.
By Resolution Rate
1
Opus 5
42.7%
2
GPT-5.6 Sol
34.6%
3
Fable 5
34.1%
View archived benchmarks

