Archived

SnorkelSequences

A procedurally-generated and expert-verified benchmark for evaluating mathematical reasoning and compositional capabilities in LLMs.
Overview
The rise of Reinforcement Learning with Verifiable Rewards (RLVR) as a paradigm in LLM post-training has given rise to a new wave of models that are highly capable in mathematics. With our new SnorkelSequences benchmark, we are not only interested in evaluating mathematical problem solving, but also the compositional capabilities of LLMs. Strong models must be able to solve unseen problems that are formed from a series of simpler tasks. This new benchmark, part of our wider series of procedurally-generated and expert-verified datasets focused on reasoning, allows for both the mathematical and compositional capabilities to be evaluated.

Leaderboard

Rank Model Score
1 gpt-5
77.6%
2 gpt-5-mini
77.6%
3 gpt-5-nano
72%
4 GPT-5.4
71.6%
5 o3-mini
71.2%
6 Gemini 2.5 Flash
70.8%
7 Claude Sonnet 4
70.4%
8 Grok 4 Fast Reasoning
70.2%
9 o4 mini
68.8%
10 NVIDIA Nemotron Super 49B v1.5
66.8%
11 Gemini 2.5 Pro
66%
12 Claude Opus 4
65.6%
13 o3
65.2%
14 Grok 4
63.2%
15 Llama 4 Maverick
62%
16 Nova Premier
51.8%
17 Llama 4 Scout
48.4%
18 Claude Sonnet 3.7
47.6%
19 Magistral Medium
47.6%
20 NVIDIA Nemotron Super 49B
44.8%
21 Nova Pro
41.2%
22 Nova Lite
40%
23 Grok 3
39.2%
24 Llama 3.3 70B
38.8%
25 Mistral Large
38.8%
26 Codestral
38.4%
27 GPT-4.1
36.8%
28 Nvidia 70B Instruct
36.4%
29 Kimi-K2-Thinking
36%
30 Llama 3.1 405B
35.2%
31 Nova Micro
33.6%
32 Qwen 3 235B
28%

Sample task

The initial version of this benchmark includes 250 complex samples, with questions covering diverse combinations of operators, ranges, conditions, and transforms. The following is an example:

Consider the integers from 7 to 362, inclusive.

First, keep only the numbers that have common logarithm (base 10) less than 3.

Of these numbers, count how many perfect squares there are.

Methodology

metric
accuracy@1 across all 250 questions.
ground truth
Verifiable answers provided programmatically alongside each question; no LLM
judge required.
task set
250 compositional sequence reasoning questions spanning multiple operator
complexity levels.
future work
Exploring code generation as an alternative path: models that cannot answer directly may succeed by writing executable programs.

Behind the benchmark

All questions in SnorkelSequences require the LLM to compute the outcome of a function or operator applied to a sequence of numbers. For example: “How many digits are there in the numbers between and including 1 and 100?” Using a procedural data generation process allows us to create problems with a verifiable ground truth answer, and parameterize question complexity. For SnorkelSequences, we control complexity through the following parameters:
01
Size of the range
Larger ranges require the LLM to reason over and track a greater number of elements in each sequence.
02
Number of intermediate components
The number of simple tasks can be controlled by introducing components such as conditions and transforms on the original sequence.
03
Operator complexity
The core of each question is an “operator” that instructs the LLM to compute a numerical value over the final sequence (e.g., “count the number of digits”). We hypothesize that different families of operators pose different challenges to LLMs.
Image

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
GPT-6.1 Sol
42.8%
3
Image
Opus 5.5
40.4%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
GPT-6.1 Sol
19.6%
2
Image
Grok 4.6
16.4%
3
Image
Grok 4.7
12.1%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
Opus 5.5
64.8%
2
Image
Sonnet 5.5
61.8%
3
Image
GPT-6 Astra
58.2%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.