SnorkelUnderwrite 2.0
SnorkelUnderwrite 2.0 evaluates agentic underwriting in a simulated commercial property and casualty insurance environment — partial information, cross-source reasoning, role-specific permissions, and stateful workflows.
The 200-task release covers answer correctness and, where required, safety, process quality, and final environment state.
At a glance
200
frontier tasks
72.5%
tasks with partial information
34.5%
tasks with 20K+ context
36%
tasks with expected state changes
Frontier performance
- Grok 4.7
- GLM 5.3
- Nemotron 3 Ultra 550B
- DeepSeek V4 Pro
- Kimi K3
- Fable 5.1
- Opus 5.5
- GPT-6 Astra
- Qwen 3.8 Max
- Muse Spark 1.3
- Gemini Flash 3.8
- Opus 5
- Grok 4.6
Key takeaways
Grok 4.7 led SnorkelUnderwrite 2.0 at 30.5%, narrowly ahead of GLM 5.3 at 30.4% and Nemotron 3 Ultra 550B at 30.0%. DeepSeek V4 Pro (28.5%) and Kimi K3 (27.9%) followed, then Fable 5.1 at 24.0% and Opus 5.5 at 22.7%, while the remaining models scored between 17.8% and 21.2%.
State-pass rates were substantially higher than overall performance and reached 74.1% for DeepSeek. However, no model passed more than 41% of answer checks, making final-answer completion the primary shared limitation.
Leaderboard
| Rank | Model | Pass@1 | Pass@5 | Cost / Trial | Cost / Task |
|---|---|---|---|---|---|
| 1 | Grok 4.7 |
30.5%
|
51.6%
|
$1.21 | $6.04 |
| 2 | GLM 5.3 |
30.4%
|
51.3%
|
$0.36 | $1.81 |
| 3 | Nemotron 3 Ultra 550B |
30%
|
54.8%
|
$0.32 | $1.62 |
| 4 | DeepSeek V4 Pro |
28.5%
|
54.5%
|
$0.61 | $3.04 |
| 5 | Kimi K3 |
27.9%
|
50.8%
|
$0.57 | $2.83 |
| 6 | Fable 5.1 |
24%
|
38.4%
|
$1.92 | $9.58 |
| 7 | Opus 5.5 |
22.7%
|
39.4%
|
$0.76 | $3.78 |
| 8 | GPT-6 Astra |
21.2%
|
35%
|
$1.71 | $8.56 |
| 9 | Qwen 3.8 Max |
20%
|
39.2%
|
$0.69 | $3.43 |
| 10 | Muse Spark 1.3 |
19.6%
|
32.7%
|
$0.33 | $1.63 |
| 11 | Gemini Flash 3.8 |
19.1%
|
35%
|
$0.58 | $2.92 |
| 12 | Opus 5 |
18.3%
|
34%
|
$0.82 | $4.12 |
| 13 | Grok 4.6 |
17.8%
|
31.3%
|
$0.66 | $3.31 |
Methodology
Evaluator
The benchmark runs in a self-contained Docker and OpenEnv environment with deterministic data fixtures, structured MCP tool calls, seeded failure behavior, stateful sessions, and complete trajectory logging.
scoring note
Five rollouts per task support frontier calibration against a reference-model panel. Results include task success and decomposed signals for answer correctness, final state, safety, process, and efficiency. Judge stability is checked through repeat scoring and human agreement sampling.
Behind the benchmark
Underwriting is a decision process under incomplete information. A case can require the agent to identify the right submission, ask for missing context, reconcile structured records with reference documents, apply line-of-business and authority rules, and make a recommendation that another professional could audit. SnorkelUnderwrite 2.0 measures whether the agent can maintain that chain from evidence to decision as the conversation and environment state develop.
Version 2.0 retains the previous iteration’s expert-verified, multi-turn underwriting identity while expanding the operating conditions around each decision. Subject-matter experts contribute and review the tasks, tools, scenarios, rubrics, and verifiers. The frontier subset emphasizes incomplete information, cross-source reconciliation, permission and policy boundaries, adversarial pressure, and controlled state changes.

