Enterprise Environments

SnorkelManufacturing

SnorkelManufacturing evaluates agents in a simulated discrete-parts plant, where engineering, production, quality, supplier, and safety decisions depend on evidence spread across records, documents, and specialized tools.

It concentrates on plant-floor decisions where engineering evidence, production pressure, and safety constraints collide.

At a glance

200

frontier tasks

33%

CAD tasks

37%

verifiable-output tasks

40.5%

tasks requiring state changes

Frontier performance

  • Grok 4.6
  • Grok 4.7
  • GLM 5.3
  • Fable 5.1
  • Muse Spark 1.3
  • Opus 5.5
  • Kimi K3
  • GPT-6 Astra
  • Opus 5
  • Qwen 3.8 Max
  • DeepSeek V4 Pro
  • Gemini Flash 3.8
  • Nemotron 3 Ultra 550B
Grok 4.6 ×
Grok 4.7 ×
GLM 5.3 ×
Fable 5.1 ×
Muse Spark 1.3 ×
Opus 5.5 ×
Kimi K3 ×
GPT-6 Astra ×
Opus 5 ×
Qwen 3.8 Max ×
DeepSeek V4 Pro ×
Gemini Flash 3.8 ×
Nemotron 3 Ultra 550B ×
Loading chart data...

Key takeaways

Grok 4.6 led SnorkelManufacturing at 16.4%, followed by Grok 4.7 at 12.1%, GLM 5.3 at 11.9%, and Fable 5.1 at 11.3%. No other model exceeded 10.5%, indicating broad difficulty across the field.

DeepSeek V4 Pro produced the highest state-pass rate at 79.9% but reached only 8.1% Pass@1. This gap shows that completing individual state changes did not reliably translate into end-to-end task success.

Leaderboard

Rank Model Pass@1 Pass@5 Cost / Trial Cost / Task
1 Grok 4.6
16.4%
37.4%
$1.98 $9.91
2 Grok 4.7
12.1%
29.9%
$2.26 $11.3
3 GLM 5.3
11.9%
34.8%
$0.92 $4.61
4 Fable 5.1
11.3%
29.5%
$5.89 $29.45
5 Muse Spark 1.3
10.5%
33.5%
$1.09 $5.46
6 Opus 5.5
10.4%
26.1%
$2.62 $13.09
7 Kimi K3
10.2%
31.9%
$2.62 $13.08
8 GPT-6 Astra
10%
23.7%
$7.39 $36.94
9 Opus 5
9.4%
23.6%
$3.12 $15.6
10 Qwen 3.8 Max
8.3%
24.4%
$2.23 $11.16
11 DeepSeek V4 Pro
8.1%
25.4%
$1.08 $5.42
12 Gemini Flash 3.8
7.1%
19.9%
$1.05 $5.25
13 Nemotron 3 Ultra 550B
5.4%
16.6%
$0.5 $2.51
Image

Want to evaluate your model against SnorkelManufacturing? Talk to our team

Methodology

Evaluator

The evaluator matches the expected result. Exact values, dates, enums, JSON, and required state transitions are checked programmatically. Technical reports and other semantic deliverables are assessed against task-specific rubrics. Answer correctness, environment state, safety, process quality, and efficiency are scored as separate signals.
timeout
Each run is evaluated under a fixed time and step budget defined by the release configuration. This keeps planning, tool use, and recovery comparable across models. The exact limits should be reported with every benchmark result.
integration

The benchmark runs in a self-contained Docker environment using the OpenEnv interaction standard. Structured JSON MCP tools expose plant records, documents, engineering artifacts, analysis functions, and governed actions. Sessions preserve state, support seeded failure injection, and capture the full trajectory of user interaction, tool calls, and observations.

scoring note

A plausible explanation is not enough. An agent can fail by grounding its decision in the wrong asset or revision, ignoring a hard safety constraint, or leaving a required update incomplete. Valid alternative tool paths are allowed. Process and efficiency diagnostics expose unnecessary calls, tool errors, and failure to anchor conclusions in evidence.

Behind the benchmark

Plant-floor work is not a lookup problem. A supplier report can conflict with measured telemetry. A revision mismatch can invalidate an otherwise reasonable action. An urgent production request can create pressure to bypass a safety or change-control rule.

SnorkelManufacturing makes these conflicts part of the evaluation. It tests whether an agent can hold uncertain causes as uncertain, identify the evidence that controls the decision, respect safety boundaries, and carry the correct conclusion into the operating environment.

The subset combines CAD reasoning, checkable outcomes, and state-changing work so labs can distinguish technical fluency from dependable manufacturing operation. Tasks cover CAD and tooling review, process and quality investigation, production planning, automation and safety checks, and work-instruction authoring — requiring agents to identify the relevant asset or revision, reconcile conflicting signals, and deliver a decision, artifact, or governed update.

More benchmarks

Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Opus 5.5
40.4%
3
Image
Fable 5.1
39.6%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Claude Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.