Archived

SnorkelFinance

A benchmark of expert-verified financial QA created from financial reports for evaluating AI agents on tool-calling and reasoning capabilities.
Overview

Enterprises are increasingly integrating LLMs in their own ecosystems, as agents, to solve real-world business tasks. While the field moves rapidly into LLM agent development and adoption, evaluating AI agents remains a challenge in dynamic interactions involving complex tools and planning.

We present a new benchmark, SnorkelFinance, for evaluating the performance of AI agents as a financial analyst. LLMs act as ReAct agents in our financial-agent simulation engine (built in Langgraph with Model Context Protocol), tasked with answering financial questions created from 10-K filings. Without access to the document, Agents must plan and use provided tools, including SQL calls, code executors, and others, to answer the financial questions accurately.

Leaderboard

Rank Model Score
1 gpt-5
81%
2 o3
81%
3 Gemini 3 Pro
80.34%
4 Claude Opus 4.1
80.3%
5 gpt-5-mini
79.3%
6 Claude Opus 4
78.3%
7 Claude Sonnet 3.7
77.9%
8 Claude Sonnet 4
76.6%
9 o4 mini
76.6%
10 Grok 4
74.04%
11 Grok 4 Fast Reasoning
73.45%
12 Kimi-K2-Thinking
71.7%
13 gpt-oss-120b
66.6%
14 Grok 3
65.86%
15 o3-mini
63.79%
16 GPT-4.1
62.7%
17 Nova Premier
62.06%
18 Gemini 2.5 Pro
60.6%
19 Gemini 2.5 Flash
53.1%
20 Qwen 3 235B
51.37%
21 gpt-5-nano
50%
22 NVIDIA Nemotron Super 49B v1.5
44%
23 Nova Pro
40.34%
24 Codestral
27.6%
25 Nova Lite
16.89%
26 Magistral Medium
16.2%
27 Nova Micro
14.48%
28 Mistral Large
13.4%

Sample task

The latest version contains 290 high-quality QA pairs, sampled over 5 industry verticals. Here is a sample of an agentic trace on a financial query.

User
For company Snorkel AI, what is the ratio of Finite-Lived Intangible Assets Net to Total Intangible Assets Net Excluding Goodwill for 2023?
Agent
To find the ratio of Finite-Lived Intangible Assets Net to Total Intangible Assets Net Excluding Goodwill for 2023, we first need to identify the table that contains this information. Let's start by calling get_descriptions to list the tables for Snorkel AI
TOOL CALL
get_table_names({
  "company_name": "Snorkel AI"
})
Tool :
['us_gaap_ScheduleOfShareBasedPaymentAwardStockOptionsValuationAssumptionsTableTextBlock', 'cb_ScheduleVITableTextBlock', 'us_gaap_ScheduleOfIntangibleAssetsAndGoodwillTableTextBlock', '...']
Agent
We can see that us_gaap_ScheduleOfIntangible... is the table that likely contains information on Intangible Assets and Goodwill. Let's get the schema for this table to confirm and find the relevant information.
TOOL CALL

{
    "function_call": {
        "name": "get_table_schema", 
        "args": {
            "company_name": "Snorkel AI",
            "table_name": "us_gaap_ScheduleOfIntangible..."
        }
    }
}

Tool :
{"description": "in_millions_of_u.s._dollars_as_of_december_31", "column_names": ["category","2024","2023"],...}
Tool :
...
Agent
The Total Intangible Assets Net Excluding Goodwill for 2023 was $6,775 million. The Finite-Lived Intangible Assets Net for 2023 was $3,267 million. The ratio is [ 3,267 / 6,775 ]. Let's calculate this.
TOOL CALL
{
    "function_call": {
        "name": "calculator", 
        "args": {
          "expression": "3267/6775"
        }
    }
}
Tool :
0.4822140221402214
Agent

FINAL ANSWER:

The ratio of Finite-Lived Intangible Assets Net to Total Intangible Assets Net Excluding Goodwill for 2023, when rounded to two decimal places, is approximately 0.48.

Methodology

JUDGE
Claude Sonnet 3.7 evaluates full execution traces, scoring both final outputs and intermediate reasoning steps.
DIMENSIONS
SQL query correctness, financial calculation accuracy, tool selection appropriateness, and overall task completion.
VALIDATION
Automated scoring calibrated against expert human annotations on a representative sample, achieving high inter-annotator agreement.
TASK COMPLEXITY

Ranges from basic financial data retrieval to multi-document analysis requiring multi-step reasoning chains.

Behind the benchmark

Our QA dataset is carefully verified by Snorkel’s network of financial experts, for realism and accuracy of the task data, on a 5-point scale, and curated for high realism and accuracy.

While closed-source models have a high, but similar performance, on standard STEM benchmarks, results on SnorkelFinance show significant differences in agent performance across models, highlighting that more work is needed to make agents reliable in complex domains and tasks.

SnorkelFinance is the first benchmark, to our knowledge, to measure performance of commonly used Enterprise models, on a Financial Agentic task. We observe several error modes: agents struggling to make complex SQL calls, hallucinations of tool arguments, and failing to correct course when tool executions fail.

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Opus 5.5
40.4%
3
Image
Fable 5.1
39.6%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
Grok 4.7
12.1%
3
Image
GLM 5.3
11.9%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
Opus 5.5
64.8%
2
Image
Sonnet 5.5
61.8%
3
Image
GPT-6 Astra
58.2%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.