Knowledge Work

WorkplaceAgents

WorkplaceAgents evaluates whether an agent can complete real-world professional tasks that require evidence synthesis, domain judgment, and a finished work product — judged by the correctness and usefulness of the resulting deliverable.

It spans 96 occupations across 19 sectors, from analytical and financial to healthcare, legal, and creative workflows.

At a glance

200

frontier tasks

19

sectors

96

occupations

19

resource formats

Frontier performance

  • Grok 4.6
  • Fable 5.1
  • Grok 4.7
  • Opus 5
  • GPT-6 Astra
  • Opus 5.5
  • Kimi K3
  • GLM 5.3
  • Gemini Flash 3.8
  • DeepSeek V4 Pro
  • Nemotron 3 Ultra 550B
Grok 4.6 ×
Fable 5.1 ×
Grok 4.7 ×
Opus 5 ×
GPT-6 Astra ×
Opus 5.5 ×
Kimi K3 ×
GLM 5.3 ×
Gemini Flash 3.8 ×
DeepSeek V4 Pro ×
Nemotron 3 Ultra 550B ×
Loading chart data...

Key takeaways

Grok 4.6 led WorkplaceAgents at 17.1%, followed by Fable 5.1 at 15.8% and Grok 4.7 at 15.2%. Opus 5 was close behind at 14.9%, with GPT-6 Astra and Opus 5.5 tied at 14.4%, and the remaining models scoring between 6.6% and 13.0%.

Cost per trial varied more than 38-fold across qualifying models, from $0.59 for Nemotron 3 Ultra 550B to $22.60 for GPT-6 Astra. Grok 4.6, the top performer, priced in the middle of the field at $7.84 per trial — well below Grok 4.7 ($12.86), Fable 5.1 ($18.65), and Opus 5 ($15.96), the models closest to it on accuracy. Opus 5.5 matched GPT-6 Astra's accuracy at a fraction of the cost: $6.35 per trial versus $22.60.

Leaderboard

Rank Model Pass@1 Pass@5 Cost / Trial Cost / Task
1 Grok 4.6
17.1%
28.6%
$7.84 $39.19
2 Fable 5.1
15.8%
30.8%
$18.65 $93.27
3 Grok 4.7
15.2%
29.4%
$12.86 $64.3
4 Opus 5
14.9%
34.7%
$15.96 $79.79
5 GPT-6 Astra
14.4%
23.7%
$22.6 $113.01
6 Opus 5.5
14.4%
26.4%
$6.35 $31.77
7 Kimi K3
13%
28.2%
$7.79 $38.96
8 GLM 5.3
12.8%
23.6%
$14.33 $71.64
9 Gemini Flash 3.8
11.2%
23.7%
$3.63 $18.15
10 DeepSeek V4 Pro
8.6%
35.2%
$1.61 $8.04
11 Nemotron 3 Ultra 550B
6.6%
18.4%
$0.59 $2.95
Image

Want to evaluate your model against WorkplaceAgents? Talk to our team

Methodology

Evaluator

Harbor evaluates completed work products using deterministic checks and rubric-based judging. Each rubric criterion maps to a substantive, independently verifiable requirement.
timeout
A run fails if the agent times out before producing a complete answer or required artifact. Each task is evaluated within a bounded execution environment.
integration

Instructions, reference files, supporting materials, and required output formats are packaged with each task. Agents work through the tools available in the task environment.

scoring note

Scoring prioritizes substantive correctness, reasoning, and work-product quality. Formatting criteria are constrained so they cannot dominate the evaluation.

Behind the benchmark

Professional work rarely ends with a short answer. It requires finding relevant evidence, applying domain judgment, resolving ambiguity, and producing an artifact that another person can use.

WorkplaceAgents evaluates that end-to-end process. The benchmark measures whether an agent can produce a correct, useful, and professionally defensible work product across diverse forms of knowledge work.

Tasks require multi-step reasoning and can produce documents, spreadsheets, presentations, code, and structured analyses across analytical, operational, technical, financial, scientific, administrative, healthcare, legal, and creative workflows.

More benchmarks

Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Opus 5.5
40.4%
3
Image
Fable 5.1
39.6%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
Grok 4.7
12.1%
3
Image
GLM 5.3
11.9%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.