Benchmarks for what frontier AI hasn't solved

Digital Work Index

An economically weighted measure of model performance across eight challenging benchmarks covering software engineering, logical reasoning and writing, and professional environments.
23.62
Fable 5.1
23.61
Opus 5.5
23.03
GPT-6 Astra
21.75
Opus 5
20.52
Grok 4.6
19.3
Grok 4.7
16.92
Gemini Flash 3.8
16.76
Kimi K3
16.45
GLM 5.3
16.29
Muse Spark 1.3
11.81
Qwen 3.8 Max
11.07
DeepSeek V4 Pro
6.95
Nemotron 3 Ultra 550B
All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Computer Use
Open Benchmarks Grants
OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Opus 5 · max
Opus 5 · max
44.33%
2
Opus 5 · high
Opus 5 · high
36.89%
3
Opus 5 · xhigh
Opus 5 · xhigh
33.33%
Learn more about OSWorld 2.0
View archived benchmarks
 Illution Back
Illution Front

Not more data. Better data.