Benchmarks for what frontier AI hasn't solved
Digital Work Index
An economically weighted measure of model performance across eight challenging benchmarks covering software engineering, logical reasoning and writing, and professional environments.
23.62
Fable 5.1
23.61
Opus 5.5
23.03
GPT-6 Astra
21.75
Opus 5
20.52
Grok 4.6
19.3
Grok 4.7
16.92
Gemini Flash 3.8
16.76
Kimi K3
16.45
GLM 5.3
16.29
Muse Spark 1.3
11.81
Qwen 3.8 Max
11.07
DeepSeek V4 Pro
6.95
Nemotron 3 Ultra 550B
All Models
Sep 22, 2026
Opus 5.5
RANK BY BENCHMARK
Agentic Coding 2.0
40.4%
#2/13
Agents’ Last Exam
38.2%
#1/43
SnorkelFinance 2.0
18.6%
#5/13
SnorkelLegal
31.4%
#7/13
SnorkelManufacturing
10.4%
#6/13
Sep 21, 2026
Grok 4.7
RANK BY BENCHMARK
Agentic Coding 2.0
31.9%
#7/13
Senior SWE-Bench
27.4%
#8/19
SnorkelFinance 2.0
20.2%
#3/13
SnorkelLegal
32.8%
#6/13
SnorkelManufacturing
12.1%
#2/13
Sep 03, 2026
GPT-6 Astra
RANK BY BENCHMARK
Agentic Coding 2.0
47.6%
#1/13
Agents’ Last Exam
34.2%
#2/43
SnorkelFinance 2.0
11.6%
#9/13
SnorkelLegal
20.9%
#13/13
SnorkelManufacturing
10%
#8/13
Sep 02, 2026
Gemini Flash 3.8
RANK BY BENCHMARK
Agentic Coding 2.0
32.6%
#6/13
SnorkelFinance 2.0
5.4%
#12/13
SnorkelLegal
25.3%
#12/13
SnorkelManufacturing
7.1%
#12/13
SnorkelRevOps
5.4%
#13/13
All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Computer Use
Open Benchmarks Grants
OSWorld 2.0
Evaluates computer-use agents on 108 long-horizon workflows.
By Binary Accuracy (500 steps)
1
Opus 5 · max
44.33%
2
Opus 5 · high
36.89%
3
Opus 5 · xhigh
33.33%
View archived benchmarks

