Archived

SnorkelGraph

A procedurally-generated and expert-verified benchmark for evaluating mathematical and spatial reasoning capabilities of LLMs through graph reasoning problems.

Overview

Many real-world applications require an LLM to perform multi-hop long-context reasoning over an input, and in the process uncover the explicit and implicit relationships that typically exist within unstructured text. With our new SnorkelGraph benchmark, we evaluate these capabilities when reasoning across more formal structures with verifiable ground truth. This new benchmark, part of our wider series of procedurally-generated and expert-verified datasets, focused on reasoning, allows for both multi-hop and mathematical reasoning capabilities to be evaluated.

Leaderboard

Rank Model Score
1 GPT-5.4
84.5%
2 Grok 4 Fast Reasoning
75%
3 o4 mini
75%
4 gpt-5-mini
72.5%
5 gpt-5
72%
6 o3
71.5%
7 o3-mini
71%
8 Claude Opus 4
64.5%
9 Grok 3
64%
10 GPT-4.1
63%
11 gpt-5-nano
62.5%
12 Qwen 3 235B
61.5%
13 Grok 4
61%
14 Claude Sonnet 4
58%
15 Gemini 2.5 Pro
58%
16 Gemini 2.5 Flash
55%
17 Magistral Medium
53.5%
18 Claude Sonnet 3.7
50%
19 Nova Premier
34.5%
20 Llama 4 Maverick
34%
21 Mistral Large
30%
22 Nvidia nemotron super 49B
29%
23 Nova Pro
28%
24 Llama 4 Scout
26%
25 Codestral
24.5%
26 Llama 3.3 70B
23.5%
27 Nvidia 70B Instruct
22.5%
28 Llama 3.1 405B
20.5%
29 Nova Lite
19%
30 Nova Micro
17.5%
31 Command R+
15%
32 Command-Light
10.5%
33 Command
10%

Sample task

The initial version of this benchmark contains 200 QA pairs, with questions covering diverse combinations of operators, ranges, and conditions. The following is an example:
Image

Find the minimum vertex cover of an undirected graph with 15 nodes. Find the smallest set of vertices that cover all edges.

Graph:
Nodes: [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14]

Edges: [(0, 1), (0, 0), (0, 2), (0, 3), (0, 4), (1, 2), (2, 2), (2, 3), (2, 5), (2, 7), (2, 8), (2, 9), (2, 10), (2, 11), (2, 14), (3, 13), (4, 6), (4, 12), (8, 8)]

Methodology

metric
accuracy@1 across all 200 questions, compared against verifiable ground-truth
answer sets.
Token Budget
16,384 tokens max (default parameters). One prompt per question; full chain-of-thought plus final answer required.
Answer Parsing

Final answers are reformatted into a canonical representation by a secondary LLM (GPT-4o), then validated by a programmatic graph validator.

Task Set
200 graph reasoning problems, including combinatorial and structural tasks such as minimum vertex cover.
note on grok-4 results
Grok-4’s accuracy is 61% under the default 16,384-token limit. At this setting, the model did not produce final answers in a significant portion of cases. With a higher 65,536-token budget, accuracy rose to 69.5%. We report the default result for comparability but note this sensitivity to token limits for context.

Behind the benchmark

All questions in the SnorkelGraph dataset require the LLM to compute the outcome of a natural language question (an operator) asked over a graph structure encoded in natural language through node and edge lists. For example: “Find the minimum density subgraph over …”. Using a procedural data generation process allows the creation of problems with a verifiable ground truth answer, which is confirmed by experts (or set of allowable answers), and parameterize question complexity. For SnorkelGraph, we control complexity through the following parameters:

01
Size of the graph 
Larger graphs (either by number of nodes or edges) require the LLM to reason over and track a greater number of elements to answer a question.
02
Operator complexity
The core of each question is an "operator" that instructs the LLM to compute a function over a graph. We hypothesize that different families of operators pose different challenges to LLMs.

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

Enterprise Environments

MBABench 2.0

Evaluates agents on building financial models in spreadsheets, across 101 expert-written tasks.

By pass rate
1
Image
Astra (Codex)
15.8%
2
Image
Fable 5.1 (Claude Code)
7.9%
3
Image
Astra (ChatGPT)
7.9%
Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
GPT-6.1 Sol
42.8%
3
Image
Opus 5.5
40.4%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
GPT-6.1 Sol
19.6%
2
Image
Grok 4.6
16.4%
3
Image
Grok 4.7
12.1%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
Opus 5.5
64.8%
2
Image
Sonnet 5.5
61.8%
3
Image
GPT-6 Astra
58.2%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.