Enterprise Environments

SnorkelFinance 2.0

SnorkelFinance 2.0 evaluates operational financial analysis in simulated investment banking and private equity workflows — incomplete requests, cross-source evidence, compliance constraints, and state changes.

The 200-task release scores the full workflow end to end: whether the conclusion is correct, supported, safe, and reflected in the environment's state.

At a glance

200

frontier tasks

93.5%

semantic-output tasks

69.5%

tasks with expected state changes

53.5%

tasks with 20K+ context

Frontier performance

  • Grok 4.6
  • GLM 5.3
  • Grok 4.7
  • Fable 5.1
  • Opus 5.5
  • Kimi K3
  • Opus 5
  • Qwen 3.8 Max
  • GPT-6 Astra
  • Muse Spark 1.3
  • DeepSeek V4 Pro
  • Gemini Flash 3.8
  • Nemotron 3 Ultra 550B
Grok 4.6 ×
GLM 5.3 ×
Grok 4.7 ×
Fable 5.1 ×
Opus 5.5 ×
Kimi K3 ×
Opus 5 ×
Qwen 3.8 Max ×
GPT-6 Astra ×
Muse Spark 1.3 ×
DeepSeek V4 Pro ×
Gemini Flash 3.8 ×
Nemotron 3 Ultra 550B ×
Loading chart data...

Key takeaways

Grok 4.6 led SnorkelFinance 2.0 on pass@1 at 25.1%, followed by GLM 5.3 at 21.2% and Grok 4.7 at 20.2%. GLM 5.3 led on pass@5 at 42.3%, ahead of Kimi K3 at 40.9%.

Every model failed more than half of the evaluated answer checks, while safety pass rates remained above 93%. This indicates that producing complete, correct answers was a greater constraint than unsafe behavior.

Leaderboard

Rank Model Pass@1 Pass@5 Cost / Trial Cost / Task
1 Grok 4.6
25.1%
37.4%
$3.46 $17.29
2 GLM 5.3
21.2%
42.3%
$0.88 $4.42
3 Grok 4.7
20.2%
35%
$2.04 $10.19
4 Fable 5.1
19.6%
34.2%
$5.83 $29.16
5 Opus 5.5
18.6%
32%
$4.01 $20.06
6 Kimi K3
17.3%
40.9%
$1.78 $8.9
7 Opus 5
16.5%
33.5%
$3.2 $16
8 Qwen 3.8 Max
14.1%
32.5%
$1.71 $8.57
9 GPT-6 Astra
11.6%
19.3%
$10.6 $53.02
10 Muse Spark 1.3
10%
24%
$0.98 $4.91
11 DeepSeek V4 Pro
9.5%
26.5%
$1.25 $6.25
12 Gemini Flash 3.8
5.4%
16%
$1.11 $5.53
13 Nemotron 3 Ultra 550B
4.3%
13.6%
$0.42 $2.09
Image

Want to evaluate your model against SnorkelFinance 2.0? Talk to our team

Methodology

Evaluator

Exact outputs and required state changes are verified programmatically. Semantic analyses are scored against task rubrics, with additional checks for evidence use, compliance, process, and efficiency. Plausible but unsupported or noncompliant answers fail.
timeout
Rollouts are bounded by the fixed execution budget in the release configuration. The time and step limits should be published with results so model comparisons remain reproducible.
integration

Agents operate inside a Docker and OpenEnv environment with structured financial data, reference materials, MCP tools, stateful sessions, seeded failure behavior, and full trace logging. The setup requires evidence gathering and tool use within the environment.

scoring note

Scenario rubrics provide the primary score, with instance-level rubrics for selected complex tasks. Five calibration rollouts per task support difficulty measurement. Report task success with financial, state, safety, process, and efficiency diagnostics.

Behind the benchmark

Financial analysis in a deal team rarely begins with a complete question or a single source of truth. An agent may need to resolve the company or transaction, reconcile market and deal data, test assumptions, check restrictions, and communicate a conclusion that others can act on. SnorkelFinance 2.0 makes these dependencies part of the measured workflow, including the state changes required to complete the task.

Version 2.0 preserves V1’s expert-verified, tool-using financial reasoning and expands it from report-grounded QA into operational investment banking and private equity workflows. Experts define and review the workflows, data, rubrics, and verifiers. The evaluation now surfaces evidence traceability, analytical judgment, compliance-aware tool use, user interaction, and correct execution in a stateful deal environment.

More benchmarks

Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Opus 5.5
40.4%
3
Image
Fable 5.1
39.6%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
Grok 4.7
12.1%
3
Image
GLM 5.3
11.9%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Claude Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.