Benchmarks for what frontier AI hasn't solved

Digital Work Index

An economically weighted measure of model performance across eight challenging benchmarks covering software engineering, logical reasoning and writing, and professional environments.
23.62
Fable 5.1
23.61
Opus 5.5
23.03
GPT-6 Astra
21.75
Opus 5
20.52
Grok 4.6
19.3
Grok 4.7
16.92
Gemini Flash 3.8
16.76
Kimi K3
16.45
GLM 5.3
16.29
Muse Spark 1.3
11.81
Qwen 3.8 Max
11.07
DeepSeek V4 Pro
6.95
Nemotron 3 Ultra 550B
All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Knowledge Work
Proprietary
WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Grok 4.6
Grok 4.6
17.1%
2
Fable 5.1
Fable 5.1
15.8%
3
Grok 4.7
Grok 4.7
15.2%
Learn more about WorkplaceAgents
Knowledge Work
Open Benchmarks Grants
Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Claude Opus 5.5
Claude Opus 5.5
38.2%
2
GPT-6 Astra
GPT-6 Astra
34.2%
3
GPT-6 Sol
GPT-6 Sol
32.2%
Learn more about Agents’ Last Exam
View archived benchmarks
 Illution Back
Illution Front

Not more data. Better data.