WorkplaceAgents
WorkplaceAgents evaluates whether an agent can complete real-world professional tasks that require evidence synthesis, domain judgment, and a finished work product — judged by the correctness and usefulness of the resulting deliverable.
It spans 96 occupations across 19 sectors, from analytical and financial to healthcare, legal, and creative workflows.
At a glance
200
frontier tasks
19
sectors
96
occupations
19
resource formats
Frontier performance
- Grok 4.6
- Fable 5.1
- Grok 4.7
- Opus 5
- GPT-6 Astra
- Opus 5.5
- Kimi K3
- GLM 5.3
- Gemini Flash 3.8
- DeepSeek V4 Pro
- Nemotron 3 Ultra 550B
Key takeaways
Grok 4.6 led WorkplaceAgents at 17.1%, followed by Fable 5.1 at 15.8% and Grok 4.7 at 15.2%. Opus 5 was close behind at 14.9%, with GPT-6 Astra and Opus 5.5 tied at 14.4%, and the remaining models scoring between 6.6% and 13.0%.
Cost per trial varied more than 38-fold across qualifying models, from $0.59 for Nemotron 3 Ultra 550B to $22.60 for GPT-6 Astra. Grok 4.6, the top performer, priced in the middle of the field at $7.84 per trial — well below Grok 4.7 ($12.86), Fable 5.1 ($18.65), and Opus 5 ($15.96), the models closest to it on accuracy. Opus 5.5 matched GPT-6 Astra's accuracy at a fraction of the cost: $6.35 per trial versus $22.60.
Leaderboard
| Rank | Model | Pass@1 | Pass@5 | Cost / Trial | Cost / Task |
|---|---|---|---|---|---|
| 1 | Grok 4.6 |
17.1%
|
28.6%
|
$7.84 | $39.19 |
| 2 | Fable 5.1 |
15.8%
|
30.8%
|
$18.65 | $93.27 |
| 3 | Grok 4.7 |
15.2%
|
29.4%
|
$12.86 | $64.3 |
| 4 | Opus 5 |
14.9%
|
34.7%
|
$15.96 | $79.79 |
| 5 | GPT-6 Astra |
14.4%
|
23.7%
|
$22.6 | $113.01 |
| 6 | Opus 5.5 |
14.4%
|
26.4%
|
$6.35 | $31.77 |
| 7 | Kimi K3 |
13%
|
28.2%
|
$7.79 | $38.96 |
| 8 | GLM 5.3 |
12.8%
|
23.6%
|
$14.33 | $71.64 |
| 9 | Gemini Flash 3.8 |
11.2%
|
23.7%
|
$3.63 | $18.15 |
| 10 | DeepSeek V4 Pro |
8.6%
|
35.2%
|
$1.61 | $8.04 |
| 11 | Nemotron 3 Ultra 550B |
6.6%
|
18.4%
|
$0.59 | $2.95 |
Want to evaluate your model against WorkplaceAgents? Talk to our team
Methodology
Evaluator
Instructions, reference files, supporting materials, and required output formats are packaged with each task. Agents work through the tools available in the task environment.
scoring note
Scoring prioritizes substantive correctness, reasoning, and work-product quality. Formatting criteria are constrained so they cannot dominate the evaluation.
Behind the benchmark
Professional work rarely ends with a short answer. It requires finding relevant evidence, applying domain judgment, resolving ambiguity, and producing an artifact that another person can use.
WorkplaceAgents evaluates that end-to-end process. The benchmark measures whether an agent can produce a correct, useful, and professionally defensible work product across diverse forms of knowledge work.
Tasks require multi-step reasoning and can produce documents, spreadsheets, presentations, code, and structured analyses across analytical, operational, technical, financial, scientific, administrative, healthcare, legal, and creative workflows.

