SnorkelFinance 2.0
SnorkelFinance 2.0 evaluates operational financial analysis in simulated investment banking and private equity workflows — incomplete requests, cross-source evidence, compliance constraints, and state changes.
The 200-task release scores the full workflow end to end: whether the conclusion is correct, supported, safe, and reflected in the environment's state.
At a glance
200
frontier tasks
93.5%
semantic-output tasks
69.5%
tasks with expected state changes
53.5%
tasks with 20K+ context
Frontier performance
- Grok 4.6
- GLM 5.3
- Grok 4.7
- Fable 5.1
- Opus 5.5
- Kimi K3
- Opus 5
- Qwen 3.8 Max
- GPT-6 Astra
- Muse Spark 1.3
- DeepSeek V4 Pro
- Gemini Flash 3.8
- Nemotron 3 Ultra 550B
Key takeaways
Grok 4.6 led SnorkelFinance 2.0 on pass@1 at 25.1%, followed by GLM 5.3 at 21.2% and Grok 4.7 at 20.2%. GLM 5.3 led on pass@5 at 42.3%, ahead of Kimi K3 at 40.9%.
Every model failed more than half of the evaluated answer checks, while safety pass rates remained above 93%. This indicates that producing complete, correct answers was a greater constraint than unsafe behavior.
Leaderboard
| Rank | Model | Pass@1 | Pass@5 | Cost / Trial | Cost / Task |
|---|---|---|---|---|---|
| 1 | Grok 4.6 |
25.1%
|
37.4%
|
$3.46 | $17.29 |
| 2 | GLM 5.3 |
21.2%
|
42.3%
|
$0.88 | $4.42 |
| 3 | Grok 4.7 |
20.2%
|
35%
|
$2.04 | $10.19 |
| 4 | Fable 5.1 |
19.6%
|
34.2%
|
$5.83 | $29.16 |
| 5 | Opus 5.5 |
18.6%
|
32%
|
$4.01 | $20.06 |
| 6 | Kimi K3 |
17.3%
|
40.9%
|
$1.78 | $8.9 |
| 7 | Opus 5 |
16.5%
|
33.5%
|
$3.2 | $16 |
| 8 | Qwen 3.8 Max |
14.1%
|
32.5%
|
$1.71 | $8.57 |
| 9 | GPT-6 Astra |
11.6%
|
19.3%
|
$10.6 | $53.02 |
| 10 | Muse Spark 1.3 |
10%
|
24%
|
$0.98 | $4.91 |
| 11 | DeepSeek V4 Pro |
9.5%
|
26.5%
|
$1.25 | $6.25 |
| 12 | Gemini Flash 3.8 |
5.4%
|
16%
|
$1.11 | $5.53 |
| 13 | Nemotron 3 Ultra 550B |
4.3%
|
13.6%
|
$0.42 | $2.09 |
Want to evaluate your model against SnorkelFinance 2.0? Talk to our team
Methodology
Evaluator
Agents operate inside a Docker and OpenEnv environment with structured financial data, reference materials, MCP tools, stateful sessions, seeded failure behavior, and full trace logging. The setup requires evidence gathering and tool use within the environment.
scoring note
Scenario rubrics provide the primary score, with instance-level rubrics for selected complex tasks. Five calibration rollouts per task support difficulty measurement. Report task success with financial, state, safety, process, and efficiency diagnostics.
Behind the benchmark
Financial analysis in a deal team rarely begins with a complete question or a single source of truth. An agent may need to resolve the company or transaction, reconcile market and deal data, test assumptions, check restrictions, and communicate a conclusion that others can act on. SnorkelFinance 2.0 makes these dependencies part of the measured workflow, including the state changes required to complete the task.
Version 2.0 preserves V1’s expert-verified, tool-using financial reasoning and expands it from report-grounded QA into operational investment banking and private equity workflows. Experts define and review the workflows, data, rubrics, and verifiers. The evaluation now surfaces evidence traceability, analytical judgment, compliance-aware tool use, user interaction, and correct execution in a stateful deal environment.

