Computer Use
Open Benchmarks Grants

OSWorld 2.0

A benchmark for evaluating computer-use agents on long-horizon, real-world workflows: 108 authentic tasks across 31 self-hosted web environments and professional desktop applications, with partial-credit scoring (avg 27.25 checkpoints per task).

Built with
Image
Image

At a glance

108

Long-horizon tasks

31

self-hosted websites

69.6%

take humans over 1 hour

>250

average agent steps

27.25

avg scoring checkpoints

Leaderboard

Rank Model Effort Approach Binary Partial Est. Cost
1 Claude Opus 5 max Batch tool
44.33%
77.67%
n/a
2 Claude Opus 5 high Batch tool
36.89%
68.92%
n/a
3 Claude Opus 5 xhigh Batch tool
33.33%
73.11%
n/a
4 Claude Opus 5 medium Batch tool
33.01%
64.85%
n/a
5 Claude Opus 5 max Batch tool
31.43%
68.31%
n/a
6 Claude Opus 5 xhigh Batch tool
30.21%
67.67%
n/a
7 Claude Opus 5 high Batch tool
28.98%
63.79%
n/a
8 GPT-5.6 Sol max Batch tool
27.34%
62.72%
n/a
9 Claude Opus 5 medium Batch tool
25.33%
60.49%
n/a
10 Claude Opus 5 low Batch tool
22.32%
54.89%
n/a
11 Claude Opus 4.8 max Batched tool
20.6%
54.8%
n/a
12 Claude Opus 5 low Batch tool
18.81%
59.09%
n/a
13 Claude Opus 4.8 max Standard
18.52%
49.33%
n/a
14 Claude Opus 4.7 max Batched tool
18.2%
48.91%
n/a
15 Claude Opus 4.7 max Standard
13.9%
49.1%
$35.8
16 GPT-5.5 xhigh Batch tool
13%
49.5%
$25.5
17 Claude Sonnet 4.6 medium Standard
9.3%
33.9%
$14.4
18 Claude Sonnet 4.6 max Standard
8.3%
41.5%
$22.3
19 MiniMax M3 enabled Standard
4.6%
22.3%
$2.4
20 Kimi K2.6 enabled Standard
4.6%
22.1%
$6.6
21 Qwen3 7-Plus thinking Standard
2.8%
21.5%
$3.8
Rank Model Effort Approach Binary Partial Est. Cost
1 GPT-5.5 xhigh Batch tool
13%
49.5%
$2.75K
2 Claude Opus 4.7 max Standard
13%
39.8%
$2.47K
3 Claude Sonnet 4.6 medium Standard
8.3%
29.4%
$990
4 Claude Sonnet 4.6 max Standard
6.5%
35.8%
$1.72K
5 Kimi K2.6 enabled Standard
4.6%
14.4%
$604
6 MiniMax M3 enabled Standard
3.7%
16.6%
$182.03
7 Qwen3 7-Plus thinking Standard
1.9%
16.6%
$403.33
Rank Model Effort Approach Binary Partial Est. Cost
1 GPT-5.5 xhigh Batch tool
13%
46.7%
$1.88K
2 Claude Opus 4.7 max Standard
4.6%
20.3%
$1.03K
3 Claude Sonnet 4.6 max Standard
4.6%
20%
$800
4 Claude Sonnet 4.6 medium Standard
4.6%
14.2%
$410
5 MiniMax M3 enabled Standard
1.9%
8.2%
$86.82
6 Kimi K2.6 enabled Standard
1.9%
7.1%
$336

OSWorld 1.0 vs 2.0

OSWorld 2.0 is a substantial expansion of the original OSWorld evaluation: tasks span far more agent steps, cross more applications, run inside reproducible self-hosted environments, and use partial credit instead of binary completion alone.

Capability

OSWorld 1.0

OSWorld 2.0

Task horizon
<30 average agent steps
>250 average agent steps
Cross-application tasks
Supported in a minority
Majority; information-dependent
Self-hosted web environments
—
31 websites
Input artifacts
Mixed / synthetic
Authentic
Challenge categories
—

10 challenge phenomena

Scoring
Binary
Partial reward; 27.25 checkpoints on average
Safety audit
—
8 diagnostic checks
User interaction
—
Simulated user

Methodology

Metrics

Binary completion (strict success) and Score (partial credit averaged across an average of 27.25 scoring checkpoints per task). Both computed against deterministic graders; 11.53% of the score is from validated model-based checks.
Self-hosted environments
31 self-hosted, high-fidelity websites with deterministic initialization, isolated execution, and reliable final-state scoring. Preserves open-web access while ensuring reproducibility across runs.
Step budgets

Submissions are scored at 150 / 300 / 500 agent-step budgets. The 500-step budget mirrors realistic long-horizon work; the 150-step budget surfaces efficiency.

Safety audit

A separate audit pipeline runs 8 diagnostic checks on each trajectory. Safety reports are scored independently from task completion.

What the results show

Three patterns recur across the evaluated agents.

1

Higher scores require disproportionately more tokens.

Crossing the 50% partial-score threshold requires order-of-magnitude more tokens than reaching 25%. Efficiency scales worse than capability.

2

Task horizon remains a hard limit.

Binary completion collapses as task length grows. On the longest workflows in the corpus, top frontier agents approach near-zero end-to-end completion regardless of step budget.

3

Agents are weak at recovering and maintaining hidden state.

When tasks require tracking unobserved or evolving context across steps, agents lose track — repeating earlier work, missing updates, or executing from stale plans.

Behind the benchmark

Computer-use agents are increasingly deployed on multi-hour professional workflows, but most benchmarks evaluate them on short desktop tasks that finish in under 30 steps. OSWorld 2.0 reframes the problem around long-horizon work, sourced from realistic end-to-end workflows that a skilled human typically takes over an hour to complete.

Acknowledgments

OSWorld 2.0 was developed by Mengqi Yuan and colleagues at XLANG Lab, University of Hong Kong. Snorkel AI is proud to support the project as a research and data partner through the Open Benchmarks Grants program. Snorkel AI contributors include Zhengyang (Jason) Qi, Vincent Sunn Chen, and Frederic Sala.

FAQs

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Opus 5.5
40.4%
3
Image
Fable 5.1
39.6%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
Grok 4.7
12.1%
3
Image
GLM 5.3
11.9%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.