SnorkelManufacturing
SnorkelManufacturing evaluates agents in a simulated discrete-parts plant, where engineering, production, quality, supplier, and safety decisions depend on evidence spread across records, documents, and specialized tools.
It concentrates on plant-floor decisions where engineering evidence, production pressure, and safety constraints collide.
At a glance
200
frontier tasks
33%
CAD tasks
37%
verifiable-output tasks
40.5%
tasks requiring state changes
Frontier performance
- Grok 4.6
- Grok 4.7
- GLM 5.3
- Fable 5.1
- Muse Spark 1.3
- Opus 5.5
- Kimi K3
- GPT-6 Astra
- Opus 5
- Qwen 3.8 Max
- DeepSeek V4 Pro
- Gemini Flash 3.8
- Nemotron 3 Ultra 550B
Key takeaways
Grok 4.6 led SnorkelManufacturing at 16.4%, followed by Grok 4.7 at 12.1%, GLM 5.3 at 11.9%, and Fable 5.1 at 11.3%. No other model exceeded 10.5%, indicating broad difficulty across the field.
DeepSeek V4 Pro produced the highest state-pass rate at 79.9% but reached only 8.1% Pass@1. This gap shows that completing individual state changes did not reliably translate into end-to-end task success.
Leaderboard
| Rank | Model | Pass@1 | Pass@5 | Cost / Trial | Cost / Task |
|---|---|---|---|---|---|
| 1 | Grok 4.6 |
16.4%
|
37.4%
|
$1.98 | $9.91 |
| 2 | Grok 4.7 |
12.1%
|
29.9%
|
$2.26 | $11.3 |
| 3 | GLM 5.3 |
11.9%
|
34.8%
|
$0.92 | $4.61 |
| 4 | Fable 5.1 |
11.3%
|
29.5%
|
$5.89 | $29.45 |
| 5 | Muse Spark 1.3 |
10.5%
|
33.5%
|
$1.09 | $5.46 |
| 6 | Opus 5.5 |
10.4%
|
26.1%
|
$2.62 | $13.09 |
| 7 | Kimi K3 |
10.2%
|
31.9%
|
$2.62 | $13.08 |
| 8 | GPT-6 Astra |
10%
|
23.7%
|
$7.39 | $36.94 |
| 9 | Opus 5 |
9.4%
|
23.6%
|
$3.12 | $15.6 |
| 10 | Qwen 3.8 Max |
8.3%
|
24.4%
|
$2.23 | $11.16 |
| 11 | DeepSeek V4 Pro |
8.1%
|
25.4%
|
$1.08 | $5.42 |
| 12 | Gemini Flash 3.8 |
7.1%
|
19.9%
|
$1.05 | $5.25 |
| 13 | Nemotron 3 Ultra 550B |
5.4%
|
16.6%
|
$0.5 | $2.51 |
Want to evaluate your model against SnorkelManufacturing? Talk to our team
Methodology
Evaluator
The benchmark runs in a self-contained Docker environment using the OpenEnv interaction standard. Structured JSON MCP tools expose plant records, documents, engineering artifacts, analysis functions, and governed actions. Sessions preserve state, support seeded failure injection, and capture the full trajectory of user interaction, tool calls, and observations.
scoring note
A plausible explanation is not enough. An agent can fail by grounding its decision in the wrong asset or revision, ignoring a hard safety constraint, or leaving a required update incomplete. Valid alternative tool paths are allowed. Process and efficiency diagnostics expose unnecessary calls, tool errors, and failure to anchor conclusions in evidence.
Behind the benchmark
Plant-floor work is not a lookup problem. A supplier report can conflict with measured telemetry. A revision mismatch can invalidate an otherwise reasonable action. An urgent production request can create pressure to bypass a safety or change-control rule.
SnorkelManufacturing makes these conflicts part of the evaluation. It tests whether an agent can hold uncertain causes as uncertain, identify the evidence that controls the decision, respect safety boundaries, and carry the correct conclusion into the operating environment.
The subset combines CAD reasoning, checkable outcomes, and state-changing work so labs can distinguish technical fluency from dependable manufacturing operation. Tasks cover CAD and tooling review, process and quality investigation, production planning, automation and safety checks, and work-instruction authoring — requiring agents to identify the relevant asset or revision, reconcile conflicting signals, and deliver a decision, artifact, or governed update.

