Enterprise Environments

SnorkelUnderwrite 2.0

SnorkelUnderwrite 2.0 evaluates agentic underwriting in a simulated commercial property and casualty insurance environment — partial information, cross-source reasoning, role-specific permissions, and stateful workflows.

The 200-task release covers answer correctness and, where required, safety, process quality, and final environment state.

At a glance

200

frontier tasks

72.5%

tasks with partial information

34.5%

tasks with 20K+ context

36%

tasks with expected state changes

Frontier performance

  • Grok 4.7
  • GLM 5.3
  • Nemotron 3 Ultra 550B
  • DeepSeek V4 Pro
  • Kimi K3
  • Fable 5.1
  • Opus 5.5
  • GPT-6 Astra
  • Qwen 3.8 Max
  • Muse Spark 1.3
  • Gemini Flash 3.8
  • Opus 5
  • Grok 4.6
Grok 4.7 ×
GLM 5.3 ×
Nemotron 3 Ultra 550B ×
DeepSeek V4 Pro ×
Kimi K3 ×
Fable 5.1 ×
Opus 5.5 ×
GPT-6 Astra ×
Qwen 3.8 Max ×
Muse Spark 1.3 ×
Gemini Flash 3.8 ×
Opus 5 ×
Grok 4.6 ×
Loading chart data...

Key takeaways

Grok 4.7 led SnorkelUnderwrite 2.0 at 30.5%, narrowly ahead of GLM 5.3 at 30.4% and Nemotron 3 Ultra 550B at 30.0%. DeepSeek V4 Pro (28.5%) and Kimi K3 (27.9%) followed, then Fable 5.1 at 24.0% and Opus 5.5 at 22.7%, while the remaining models scored between 17.8% and 21.2%.

State-pass rates were substantially higher than overall performance and reached 74.1% for DeepSeek. However, no model passed more than 41% of answer checks, making final-answer completion the primary shared limitation.

Leaderboard

Rank Model Pass@1 Pass@5 Cost / Trial Cost / Task
1 Grok 4.7
30.5%
51.6%
$1.21 $6.04
2 GLM 5.3
30.4%
51.3%
$0.36 $1.81
3 Nemotron 3 Ultra 550B
30%
54.8%
$0.32 $1.62
4 DeepSeek V4 Pro
28.5%
54.5%
$0.61 $3.04
5 Kimi K3
27.9%
50.8%
$0.57 $2.83
6 Fable 5.1
24%
38.4%
$1.92 $9.58
7 Opus 5.5
22.7%
39.4%
$0.76 $3.78
8 GPT-6 Astra
21.2%
35%
$1.71 $8.56
9 Qwen 3.8 Max
20%
39.2%
$0.69 $3.43
10 Muse Spark 1.3
19.6%
32.7%
$0.33 $1.63
11 Gemini Flash 3.8
19.1%
35%
$0.58 $2.92
12 Opus 5
18.3%
34%
$0.82 $4.12
13 Grok 4.6
17.8%
31.3%
$0.66 $3.31

Methodology

Evaluator

Task-specific reward modules combine deterministic checks for exact outputs and final environment state with semantic judging for meaning-based responses. Safety and process checks assess policy compliance, tool use, exceptions, redundant actions, and efficiency.
timeout
Runs use fixed time and step limits defined by the evaluation configuration so models are compared under the same operating conditions.
integration

The benchmark runs in a self-contained Docker and OpenEnv environment with deterministic data fixtures, structured MCP tool calls, seeded failure behavior, stateful sessions, and complete trajectory logging.

scoring note

Five rollouts per task support frontier calibration against a reference-model panel. Results include task success and decomposed signals for answer correctness, final state, safety, process, and efficiency. Judge stability is checked through repeat scoring and human agreement sampling.

Behind the benchmark

Underwriting is a decision process under incomplete information. A case can require the agent to identify the right submission, ask for missing context, reconcile structured records with reference documents, apply line-of-business and authority rules, and make a recommendation that another professional could audit. SnorkelUnderwrite 2.0 measures whether the agent can maintain that chain from evidence to decision as the conversation and environment state develop.

Version 2.0 retains the previous iteration’s expert-verified, multi-turn underwriting identity while expanding the operating conditions around each decision. Subject-matter experts contribute and review the tasks, tools, scenarios, rubrics, and verifiers. The frontier subset emphasizes incomplete information, cross-source reconciliation, permission and policy boundaries, adversarial pressure, and controlled state changes.

More benchmarks

Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Opus 5.5
40.4%
3
Image
Fable 5.1
39.6%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
Grok 4.7
12.1%
3
Image
GLM 5.3
11.9%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Claude Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.