Benchmarks for what frontier AI hasn't solved
Digital Work Index
An economically weighted measure of model performance across eight challenging benchmarks covering software engineering, logical reasoning and writing, and professional environments.
23.62
Fable 5.1
23.61
Opus 5.5
23.03
GPT-6 Astra
21.75
Opus 5
20.52
Grok 4.6
19.3
Grok 4.7
16.92
Gemini Flash 3.8
16.76
Kimi K3
16.45
GLM 5.3
16.29
Muse Spark 1.3
11.81
Qwen 3.8 Max
11.07
DeepSeek V4 Pro
6.95
Nemotron 3 Ultra 550B
All Models
Sep 22, 2026
Opus 5.5
RANK BY BENCHMARK
Agentic Coding 2.0
40.4%
#2/13
Agents’ Last Exam
38.2%
#1/43
SnorkelFinance 2.0
18.6%
#5/13
SnorkelLegal
31.4%
#7/13
SnorkelManufacturing
10.4%
#6/13
Sep 21, 2026
Grok 4.7
RANK BY BENCHMARK
Agentic Coding 2.0
31.9%
#7/13
Senior SWE-Bench
27.4%
#8/19
SnorkelFinance 2.0
20.2%
#3/13
SnorkelLegal
32.8%
#6/13
SnorkelManufacturing
12.1%
#2/13
Sep 03, 2026
GPT-6 Astra
RANK BY BENCHMARK
Agentic Coding 2.0
47.6%
#1/13
Agents’ Last Exam
34.2%
#2/43
SnorkelFinance 2.0
11.6%
#9/13
SnorkelLegal
20.9%
#13/13
SnorkelManufacturing
10%
#8/13
Sep 02, 2026
Gemini Flash 3.8
RANK BY BENCHMARK
Agentic Coding 2.0
32.6%
#6/13
SnorkelFinance 2.0
5.4%
#12/13
SnorkelLegal
25.3%
#12/13
SnorkelManufacturing
7.1%
#12/13
SnorkelRevOps
5.4%
#13/13
All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Science & Research
Open Benchmarks Grants
Terminal-Bench-Science
Evaluates agents on scientific workflows derived from researchers’ own work.
By Resolution Rate
1
GPT-6 Astra
68.1%
2
Opus 5.5
63.3%
3
Fable 5.1
40.0%
View archived benchmarks

