Benchmarks for what frontier AI hasn't solved
Digital Work Index
Evaluates coding agents on senior-level software engineering tasks.
Measures how well AI agents design libraries that other agents can use.
Evaluates agents on scientific workflows derived from researchers’ own work.
Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.
Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.
Evaluates terminal agents on continuously updated software-engineering tasks.
Evaluates validated multi-file code changes in open-source repositories.
Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.
Evaluates computer-use agents on 108 long-horizon workflows.
Evaluates revenue-operations workflows involving systems, controls, and business-state updates.
Evaluates financial workflows requiring evidence gathering, analysis, and compliance.
Measures coding-agent performance across evolving software requirements.
Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.
Measures improvement across sequential, stateful tasks.
Evaluates legal workflows requiring evidence, procedure, authority, and action.
Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.
Evaluates terminal agents on containerized tasks across seven domains.

