Digital Work Index

The Digital Work Index is an economically weighted measure of model performance across eight, frontier-difficulty benchmarks in software engineering, logical reasoning and writing, and professional environments.

Leaderboard

Rank Model Score
1 Fable 5.1
23.62
2 Opus 5.5
23.61
3 GPT-6 Astra
23.03
4 Opus 5
21.75
5 Grok 4.6
20.52
6 Grok 4.7
19.3
7 Gemini Flash 3.8
16.92
8 Kimi K3
16.76
9 GLM 5.3
16.45
10 Muse Spark 1.3
16.29
11 Qwen 3.8 Max
11.81
12 DeepSeek V4 Pro
11.07
13 Nemotron 3 Ultra 550B
6.95
Rank Model Score
1 GPT-6 Astra
28.76
2 Opus 5.5
27.66
3 Fable 5.1
26.16
4 Opus 5
25.57
5 Grok 4.6
18.67
6 Gemini Flash 3.8
18.52
7 Grok 4.7
18.09
8 Kimi K3
14.32
9 GLM 5.3
13.44
10 Muse Spark 1.3
12.8
11 Qwen 3.8 Max
9.45
12 DeepSeek V4 Pro
8.11
13 Nemotron 3 Ultra 550B
2.64

Methodology

Scope

We began with 825 BLS occupations and retain 237 whose work primarily produces digital outputs such as documents, datasets, software, designs, or decision records. Employment multiplied by BLS May 2025 mean annual wage defines the $5.105 trillion denominator out of $10.8 trillion tracked by BLS. The eight benchmarks map to 29 occupations.
Task mapping
Every evaluated task is assigned to an occupation. Seven benchmark mappings come from a comparison of task-level research to O*NET occupational taxonomies. Harvey LAB is mapped from its public task distribution.
Adjustments and weights

Each benchmark is adjusted by the share of its mapped occupations employed in the simulated industry, using BLS industry-by-occupation data. Coding is treated as portable across industries. Specialized benchmarks are limited to their modeled industries. Shared occupations are split according to task concentration to avoid double counting. Each benchmark receives a wage term and a task-volume term. Its final weight is the normalized geometric mean of the two, giving equal influence to economic coverage and evaluated task volume.

Scoring and precision

Each model's score is the weighted average of its benchmark pass rates. With approximately 200 tasks per benchmark, estimated sampling precision is ±3.5 points per benchmark, giving a standard error of ±1.7 points on the index. Gaps between models smaller than roughly 2.4 points should not be treated as demonstrated ranking differences.

Behind the benchmark

The Digital Work Index summarizes model performance across eight challenging benchmarks at frontier difficulty covering software engineering, legal work, finance, insurance, manufacturing, and revenue operations. It combines results from Agentic Coding 2.0, SWE-bench CLI, SnorkelFinance 2.0, SnorkelUnderwrite 2.0, SnorkelManufacturing, SnorkelRevOps, SnorkelLegal, and Harvey LAB.

Each benchmark is weighted using two signals. The first is the wage base of the occupations it represents. The second is how much real-life work the benchmark reaches — measured by the jobs its tasks test. The current release covers 13 models and 29 occupations. Together, the benchmarks represent 32.7% of the estimated $5.1 trillion annual wage bill across the 237 U.S. occupations classified here as digital work. The index provides an economically weighted comparison of model capability at the frontier. Productivity, automation, and labor-market impact are outside its scope.

More benchmarks

Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Opus 5.5
40.4%
3
Image
Fable 5.1
39.6%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
Grok 4.7
12.1%
3
Image
GLM 5.3
11.9%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.