Benchmarks for what frontier AI hasn't solved

Digital Work Index

An economically weighted measure of model performance across eight challenging benchmarks covering software engineering, logical reasoning and writing, and professional environments.
23.62
Fable 5.1
23.61
Opus 5.5
23.03
GPT-6 Astra
21.75
Opus 5
20.52
Grok 4.6
19.3
Grok 4.7
16.92
Gemini Flash 3.8
16.76
Kimi K3
16.45
GLM 5.3
16.29
Muse Spark 1.3
11.81
Qwen 3.8 Max
11.07
DeepSeek V4 Pro
6.95
Nemotron 3 Ultra 550B
All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Enterprise Environments
Proprietary
SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Grok 4.6
Grok 4.6
16.4%
2
Grok 4.7
Grok 4.7
12.1%
3
GLM 5.3
GLM 5.3
11.9%
Learn more about SnorkelManufacturing
Enterprise Environments
Proprietary
SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Grok 4.6
Grok 4.6
15.8%
2
Fable 5.1
Fable 5.1
14.0%
3
Kimi K3
Kimi K3
13.1%
Learn more about SnorkelRevOps
Enterprise Environments
Proprietary
SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Grok 4.6
Grok 4.6
25.1%
2
GLM 5.3
GLM 5.3
21.2%
3
Grok 4.7
Grok 4.7
20.2%
Learn more about SnorkelFinance 2.0
Enterprise Environments
Proprietary
SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Grok 4.7
Grok 4.7
30.5%
2
GLM 5.3
GLM 5.3
30.4%
3
Nemotron 3 Ultra 550B
Nemotron 3 Ultra 550B
30.0%
Learn more about SnorkelUnderwrite 2.0
Enterprise Environments
Proprietary
SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Grok 4.6
Grok 4.6
40.7%
2
GLM 5.3
GLM 5.3
37.8%
3
Muse Spark 1.3
Muse Spark 1.3
36.5%
Learn more about SnorkelLegal
View archived benchmarks
 Illution Back
Illution Front

Not more data. Better data.