Archived

SnorkelUnderwrite

An expert-verified frontier benchmark with multi-turn conversations, focused on agentic reasoning and tool use in commercial underwriting settings.
Overview

To date, most agentic and reasoning benchmarks have centered around tasks in the STEM domains. However, in the business world, many problems require domain-specific reasoning and behavior in messier ecosystems that require logic with metadata and tool combinations. To address this gap, we have developed a set of tasks that challenges AI agents in real-world ways, using a complex enterprise domain: commercial insurance underwriting.

The SnorkelUnderwrite benchmark is multi-turn, requiring AI agents to effectively interact with underwriters to help them solve their tasks by not only reasoning over tools, but by also asking the underwriters informative questions.

Leaderboard

Rank Model Score
1 GPT-5.4
91%
2 Claude Opus 4.1
86.3%
3 gpt-5
83.33%
4 Grok 4
83.3%
5 Grok 4 Fast Reasoning
81.33%
6 Grok 3
78%
7 o4 mini
78%
8 Claude Opus 4
77%
9 o3
77%
10 Claude Sonnet 3.7
74.6%
11 Claude Sonnet 4
72.3%
12 gpt-5-mini
71.67%
13 Kimi-K2-Thinking
71.3%
14 GPT-4.1
70.6%
15 Gemini 2.5 Flash
61%
16 Nova Premier
57%
17 Gemini 2.5 Pro
56.3%
18 Nova Pro
52.3%
19 gpt-5-nano
47%
20 Llama 3.3 70B
46.3%
21 Llama 4 Maverick
46.3%
22 Llama 4 Scout
44.3%
23 o3-mini
44.3%
24 Nova Lite
40%
25 Mistral Large
38.3%
26 Codestral
34%
27 Nova Micro
31%
28 gpt-oss-120b
30%
29 Magistral Medium
29.3%
30 Command R+
25.7%
31 Qwen 3 235B
21.3%
32 Llama 3.1 405B
20%
33 Command R
15.3%

Data sample

The trace below shows a sample conversation from the benchmark:
User
Could you guide me toward the correct six-digit NAICS identifier for this company?
Agent
I'll help you find the correct NAICS (North American Industry Classification System) code for the company. To do this effectively, I need some information about the company's primary business activities. What is the company's main business or industry?
User
Managing and holding securities of startups.
Agent
I'll help you find the correct NAICS code for a company that manages and holds securities of startups. Let me look at the available tables to find relevant information about NAICS codes.
TOOL USE

{
    "id": "call_E4nhv3uQks1Y9NxXJ0LbcXK",
    "input": {},
    "name": "list_tables",
    "type": "tool_call"
}

Tool :
["appetite_guide", "lobs", "naics", "naics_2012_2017", "naics_2022_2017", "naics_2digit", "small_business_definition", "states", "supply_chain_emissions_statistics"]
Agent
Now let me search for NAICS codes related to managing and holding securities or startups.
TOOL CALL

{
    "id": "call_GiY7wnOMqrcvdRK54X7fHn2",
    "input": {
        "table_name": "naics"
    },
    "name": "get_table_schema",
    ...
}

Tool :
[{'cid': 0, 'name': 'index', 'type': 'INTEGER', 'notnull': 0, 'dflt_value': null, 'pk': 0}, {'cid': 1, 'name': 'Code', 'type': 'INTEGER', 'notnull': 0, 'dflt_value': null, 'pk': 0}, ...]
Agent
Now let me search for NAICS codes related to managing and holding securities of startups.
TOOL USE

{
    "id": "call_o1IyWlKrkxYNCAC2yLqWSzg",
   "input": {
     "query": (
         "SELECT Code, Title, Description "
         "FROM naics "
         "WHERE Title LIKE '%securities%' "
         "OR Title LIKE '%holding%' "
         "OR Title LIKE '%investment%' "
         "OR Description LIKE '%securities%' "
         "OR Description LIKE '%holding%' "
         "OR Description LIKE '%startup%' "
         "ORDER BY Code"
      )
   },
   "name": "read_query",
   "type": "tool_call"
}

Agent

FINAL ANSWER:
523910. The reason is that the company is primarily engaged in managing and holding securities of startups, which aligns with the NAICS code 523910 for “Miscellaneous Intermediation.”…

Dataset

The evaluation dataset is publicly available on Hugging Face. It covers 6 distinct task types related to
small business insurance underwriting, with multi-turn conversational traces grounded in realistic
underwriter workflows.

Methodology

metric
Overall accuracy via LLM-as-a-Judge (GPT 4.1), comparing the agent’s final answer against a programmatically generated reference.
judge agreement
94.5% agreement on a balanced random sample of 200 conversations, validated against human annotations in Snorkel Evaluate.
scope
All traces are scored, including those where agents failed to produce a final answer due to recursion errors or premature termination.
failure rate
Agent failures observed at <1% for closed-source models and 10–30%+ for open-source models.

Behind the benchmark

We built the system in LangGraph with Model Context Protocol (MCP) and ReAct Agents. We engaged with our network of Chartered Property Casualty Underwriters (CPCUs) to create crucial components of the system, with a diverse sample dataset covering 6 distinct types of tasks, all related to applications for insurance by small businesses. Many of the tasks include subtasks involving more nuanced, complex underwriting logic. In each conversation, the underwriter has one of these specific tasks to solve. The tasks require an average of 3–7 steps of reasoning and tool use, with a total of 10–20 conversational turns.

From the blog

Image for Building the benchmark: inside our agentic insurance underwriting dataset

Building the benchmark: inside our agentic insurance underwriting dataset

In this post, we unpack how Snorkel built a realistic benchmark dataset to evaluate AI agents in commercial insurance underwriting....

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
GPT-6.1 Sol
42.8%
3
Image
Opus 5.5
40.4%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
GPT-6.1 Sol
19.6%
2
Image
Grok 4.6
16.4%
3
Image
Grok 4.7
12.1%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
Opus 5.5
64.8%
2
Image
Sonnet 5.5
61.8%
3
Image
GPT-6 Astra
58.2%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.