Benchmarks for what frontier AI hasn't solved
Digital Work Index
An economically weighted measure of model performance across eight challenging benchmarks covering software engineering, logical reasoning and writing, and professional environments.
23.62
Fable 5.1
23.61
Opus 5.5
23.03
GPT-6 Astra
21.75
Opus 5
20.52
Grok 4.6
19.3
Grok 4.7
16.92
Gemini Flash 3.8
16.76
Kimi K3
16.45
GLM 5.3
16.29
Muse Spark 1.3
11.81
Qwen 3.8 Max
11.07
DeepSeek V4 Pro
6.95
Nemotron 3 Ultra 550B
All Models
Sep 22, 2026
Opus 5.5
RANK BY BENCHMARK
Agentic Coding 2.0
40.4%
#2/13
Agents’ Last Exam
38.2%
#1/43
SnorkelFinance 2.0
18.6%
#5/13
SnorkelLegal
31.4%
#7/13
SnorkelManufacturing
10.4%
#6/13
Sep 21, 2026
Grok 4.7
RANK BY BENCHMARK
Agentic Coding 2.0
31.9%
#7/13
Senior SWE-Bench
27.4%
#8/19
SnorkelFinance 2.0
20.2%
#3/13
SnorkelLegal
32.8%
#6/13
SnorkelManufacturing
12.1%
#2/13
Sep 03, 2026
GPT-6 Astra
RANK BY BENCHMARK
Agentic Coding 2.0
47.6%
#1/13
Agents’ Last Exam
34.2%
#2/43
SnorkelFinance 2.0
11.6%
#9/13
SnorkelLegal
20.9%
#13/13
SnorkelManufacturing
10%
#8/13
Sep 02, 2026
Gemini Flash 3.8
RANK BY BENCHMARK
Agentic Coding 2.0
32.6%
#6/13
SnorkelFinance 2.0
5.4%
#12/13
SnorkelLegal
25.3%
#12/13
SnorkelManufacturing
7.1%
#12/13
SnorkelRevOps
5.4%
#13/13
All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Knowledge Work
Proprietary
WorkplaceAgents
Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.
By pass@1
1
Grok 4.6
17.1%
2
Fable 5.1
15.8%
3
Grok 4.7
15.2%
Knowledge Work
Open Benchmarks Grants
Agents’ Last Exam
Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.
By Pass Rate
1
Claude Opus 5.5
38.2%
2
GPT-6 Astra
34.2%
3
GPT-6 Sol
32.2%
View archived benchmarks

