Enterprise Environments
Open Benchmarks Grants

MBABench 2.0

MBABench 2.0 evaluates whether AI agents can build financial models in spreadsheets the way a finance professional would, across 101 tasks written from scratch by finance experts.

A task passes only if every numerical answer is right and the workbook meets a professional quality bar. MBABench is built by researchers at Columbia Business School and supported through Snorkel’s Open Benchmarks Grants.

At a glance

101

tasks written by finance experts

6

finance areas covered

16

financial model types

20

agents evaluated

Key takeaways

Astra in Codex led MBABench 2.0 at 15.8%, followed by Fable 5.1 in Claude Code and Astra in ChatGPT, both at 7.9%. Ten other agents scored between 1% and 8%, and nine agents scored 0%, including Sol in ChatGPT for Excel. Only Astra in Codex passed any of the 20 Hard tasks, at 5.0%.

The harness mattered as much as the model. Astra passed 15.8% of tasks in Codex but 7.9% in ChatGPT, and Fable 5.1 passed 7.9% in Claude Code but 2.0% in Claude for Excel. Getting the numbers right was the main obstacle: the most accurate agent, Fable 5.1 in Claude Code, got every answer right on 58.9% of tasks.

Leaderboard

Rank Agent Harness Effort Pass rate ± SE Numerical accuracy ± SE Passed
1 Astra Codex xhigh
15.8%
± 3.6
54.4%
± 5.3 16 / 101
2 Fable 5.1 Claude Code max
7.9%
± 2.7
58.9%
± 5.2 8 / 101
3 Astra ChatGPT ultra
7.9%
± 2.7
51.1%
± 5.3 8 / 101
4 Fable 5.1 Claude Cowork max
5%
± 2.2
56.7%
± 5.3 5 / 101
5 Opus 5 Claude Code max
4%
± 1.9
46.7%
± 5.3 4 / 101
6 GPT-6-Pro ChatGPT n/a
3%
± 1.7
48.9%
± 5.3 3 / 101
7 Fable 5.1 Claude for Excel n/a
2%
± 1.4
52.2%
± 5.3 2 / 101
8 Opus 5 Claude for Excel n/a
1%
± 1
54.4%
± 5.3 1 / 101
9 Opus 5 Claude Cowork max
1%
± 1
50%
± 5.3 1 / 101
10 Grok 4.6 Codex xhigh
1%
± 1
45.6%
± 5.3 1 / 101
11 Qwen 3.8 Codex xhigh
1%
± 1
40%
± 5.2 1 / 101
12 Fable 5.1 In-house harness max
0%
± 0
50%
± 5.3 0 / 101
13 Sol ChatGPT for Excel xhigh
0%
± 0
45.6%
± 5.3 0 / 101
14 Gemini 3.8 Codex high
0%
± 0
45.6%
± 5.3 0 / 101
15 Kimi K3 Codex max
0%
± 0
44.4%
± 5.3 0 / 101
16 Astra In-house harness xhigh
0%
± 0
41.1%
± 5.2 0 / 101
17 Kimi K3 In-house harness max
0%
± 0
36.7%
± 5.1 0 / 101
18 Qwen 3.8 In-house harness xhigh
0%
± 0
28.9%
± 4.8 0 / 101
19 Gemini 3.8 In-house harness high
0%
± 0
15.6%
± 3.8 0 / 101
20 Grok 4.6 In-house harness xhigh
0%
± 0
7.8%
± 2.8 0 / 101
Rank Agent Harness Effort Pass rate ± SE Numerical accuracy ± SE Passed
1 Astra ChatGPT ultra
54.5%
± 15
90.9%
± 8.7 6 / 11
2 Fable 5.1 Claude Code max
45.5%
± 15
81.8%
± 11.6 5 / 11
3 Astra Codex xhigh
45.5%
± 15
72.7%
± 13.4 5 / 11
4 Opus 5 Claude Code max
18.2%
± 11.6
72.7%
± 13.4 2 / 11
5 GPT-6-Pro ChatGPT n/a
18.2%
± 11.6
72.7%
± 13.4 2 / 11
6 Fable 5.1 Claude Cowork max
9.1%
± 8.7
90.9%
± 8.7 1 / 11
7 Fable 5.1 Claude for Excel n/a
9.1%
± 8.7
81.8%
± 11.6 1 / 11
8 Opus 5 Claude for Excel n/a
9.1%
± 8.7
72.7%
± 13.4 1 / 11
9 Grok 4.6 Codex xhigh
0%
± 0
90.9%
± 8.7 0 / 11
10 Fable 5.1 In-house harness max
0%
± 0
81.8%
± 11.6 0 / 11
11 Opus 5 Claude Cowork max
0%
± 0
72.7%
± 13.4 0 / 11
12 Sol ChatGPT for Excel xhigh
0%
± 0
72.7%
± 13.4 0 / 11
13 Astra In-house harness xhigh
0%
± 0
72.7%
± 13.4 0 / 11
14 Qwen 3.8 Codex xhigh
0%
± 0
63.6%
± 14.5 0 / 11
15 Gemini 3.8 Codex high
0%
± 0
63.6%
± 14.5 0 / 11
16 Kimi K3 In-house harness max
0%
± 0
54.5%
± 15 0 / 11
17 Grok 4.6 In-house harness xhigh
0%
± 0
54.5%
± 15 0 / 11
18 Kimi K3 Codex max
0%
± 0
45.5%
± 15 0 / 11
19 Qwen 3.8 In-house harness xhigh
0%
± 0
45.5%
± 15 0 / 11
20 Gemini 3.8 In-house harness high
0%
± 0
36.4%
± 14.5 0 / 11
Rank Agent Harness Effort Pass rate ± SE Numerical accuracy ± SE Passed
1 Astra Codex xhigh
22.5%
± 6.6
55%
± 7.9 9 / 40
2 Fable 5.1 Claude Cowork max
10%
± 4.7
60%
± 7.7 4 / 40
3 Fable 5.1 Claude Code max
5%
± 3.4
62.5%
± 7.7 2 / 40
4 Astra ChatGPT ultra
5%
± 3.4
52.5%
± 7.9 2 / 40
5 Opus 5 Claude Code max
5%
± 3.4
52.5%
± 7.9 2 / 40
6 Grok 4.6 Codex xhigh
2.5%
± 2.5
57.5%
± 7.8 1 / 40
7 Fable 5.1 Claude for Excel n/a
2.5%
± 2.5
55%
± 7.9 1 / 40
8 GPT-6-Pro ChatGPT n/a
2.5%
± 2.5
50%
± 7.9 1 / 40
9 Qwen 3.8 Codex xhigh
2.5%
± 2.5
42.5%
± 7.8 1 / 40
10 Fable 5.1 In-house harness max
0%
± 0
62.5%
± 7.7 0 / 40
11 Opus 5 Claude for Excel n/a
0%
± 0
57.5%
± 7.8 0 / 40
12 Opus 5 Claude Cowork max
0%
± 0
52.5%
± 7.9 0 / 40
13 Kimi K3 Codex max
0%
± 0
52.5%
± 7.9 0 / 40
14 Astra In-house harness xhigh
0%
± 0
52.5%
± 7.9 0 / 40
15 Gemini 3.8 Codex high
0%
± 0
52.5%
± 7.9 0 / 40
16 Sol ChatGPT for Excel xhigh
0%
± 0
47.5%
± 7.9 0 / 40
17 Kimi K3 In-house harness max
0%
± 0
45%
± 7.9 0 / 40
18 Qwen 3.8 In-house harness xhigh
0%
± 0
40%
± 7.7 0 / 40
19 Gemini 3.8 In-house harness high
0%
± 0
22.5%
± 6.6 0 / 40
20 Grok 4.6 In-house harness xhigh
0%
± 0
10%
± 4.7 0 / 40
Rank Agent Harness Effort Pass rate ± SE Numerical accuracy ± SE Passed
1 Fable 5.1 Claude Code max
3.3%
± 3.3
63.3%
± 8.8 1 / 30
2 Astra Codex xhigh
3.3%
± 3.3
60%
± 8.9 1 / 30
3 Opus 5 Claude Cowork max
3.3%
± 3.3
53.3%
± 9.1 1 / 30
4 Fable 5.1 Claude Cowork max
0%
± 0
63.3%
± 8.8 0 / 30
5 Opus 5 Claude for Excel n/a
0%
± 0
60%
± 8.9 0 / 30
6 Astra ChatGPT ultra
0%
± 0
56.7%
± 9 0 / 30
7 GPT-6-Pro ChatGPT n/a
0%
± 0
56.7%
± 9 0 / 30
8 Fable 5.1 Claude for Excel n/a
0%
± 0
56.7%
± 9 0 / 30
9 Opus 5 Claude Code max
0%
± 0
53.3%
± 9.1 0 / 30
10 Sol ChatGPT for Excel xhigh
0%
± 0
53.3%
± 9.1 0 / 30
11 Fable 5.1 In-house harness max
0%
± 0
50%
± 9.1 0 / 30
12 Astra In-house harness xhigh
0%
± 0
50%
± 9.1 0 / 30
13 Gemini 3.8 Codex high
0%
± 0
50%
± 9.1 0 / 30
14 Grok 4.6 Codex xhigh
0%
± 0
46.7%
± 9.1 0 / 30
15 Kimi K3 In-house harness max
0%
± 0
43.3%
± 9 0 / 30
16 Qwen 3.8 Codex xhigh
0%
± 0
40%
± 8.9 0 / 30
17 Kimi K3 Codex max
0%
± 0
40%
± 8.9 0 / 30
18 Qwen 3.8 In-house harness xhigh
0%
± 0
30%
± 8.4 0 / 30
19 Gemini 3.8 In-house harness high
0%
± 0
16.7%
± 6.8 0 / 30
20 Grok 4.6 In-house harness xhigh
0%
± 0
10%
± 5.5 0 / 30
Rank Agent Harness Effort Pass rate ± SE Numerical accuracy ± SE Passed
1 Astra Codex xhigh
5%
± 4.9
45%
± 11.1 1 / 20
2 Fable 5.1 Claude Code max
0%
± 0
45%
± 11.1 0 / 20
3 Astra ChatGPT ultra
0%
± 0
40%
± 11 0 / 20
4 Fable 5.1 Claude Cowork max
0%
± 0
40%
± 11 0 / 20
5 Fable 5.1 Claude for Excel n/a
0%
± 0
40%
± 11 0 / 20
6 Opus 5 Claude Cowork max
0%
± 0
40%
± 11 0 / 20
7 Opus 5 Claude for Excel n/a
0%
± 0
40%
± 11 0 / 20
8 GPT-6-Pro ChatGPT n/a
0%
± 0
35%
± 10.7 0 / 20
9 Qwen 3.8 Codex xhigh
0%
± 0
35%
± 10.7 0 / 20
10 Kimi K3 Codex max
0%
± 0
35%
± 10.7 0 / 20
11 Sol ChatGPT for Excel xhigh
0%
± 0
30%
± 10.2 0 / 20
12 Opus 5 Claude Code max
0%
± 0
25%
± 9.7 0 / 20
13 Fable 5.1 In-house harness max
0%
± 0
25%
± 9.7 0 / 20
14 Gemini 3.8 Codex high
0%
± 0
25%
± 9.7 0 / 20
15 Grok 4.6 Codex xhigh
0%
± 0
20%
± 8.9 0 / 20
16 Kimi K3 In-house harness max
0%
± 0
10%
± 6.7 0 / 20
17 Astra In-house harness xhigh
0%
± 0
5%
± 4.9 0 / 20
18 Qwen 3.8 In-house harness xhigh
0%
± 0
5%
± 4.9 0 / 20
19 Grok 4.6 In-house harness xhigh
0%
± 0
0%
± 0 0 / 20
20 Gemini 3.8 In-house harness high
0%
± 0
0%
± 0 0 / 20

Pass: every numerical answer is right and the workbook has at most 1, 2, 3 and 2 other criteria flagged across the four importance tiers. Accuracy: every answer matches the key; Overall accuracy is measured on the 90 tasks rated Medium or harder. ± is one standard error.

Methodology

Evaluator

Numerical answers are checked programmatically against expert reference values. The workbook must also meet a professional quality bar before a task counts as passed.
Difficulty
Tasks are split into Easy, Medium, Medium Hard and Hard tiers, and results are reported for each. Only one agent passed any of the 20 Hard tasks.
Harnesses

Each agent is run in its own harness, such as Codex, Claude Code, ChatGPT or Claude for Excel. The same model can score very differently depending on the harness, so the leaderboard lists each pairing separately.

scoring note

Pass rate is strict. A task counts only when every numerical answer is right and the workbook meets the quality bar, so an agent can be mostly accurate and still score 0%.

Behind the benchmark

Financial modeling is more than arithmetic. A correct model needs the right structure, linked formulas and assumptions that hold together across the workbook, so one wrong cell can change every answer downstream.

MBABench 2.0 was built by researchers at Columbia Business School, led by Hongseok Namkoong and Thomson Yen. Its 101 tasks were written from scratch by finance experts and span 6 finance areas and 16 financial model types.

MBABench is supported through Snorkel’s Open Benchmarks Grants. Tasks, data and code are available on the MBABench website and GitHub.

More benchmarks

Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
GPT-6.1 Sol
42.8%
3
Image
Opus 5.5
40.4%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
GPT-6.1 Sol
19.6%
2
Image
Grok 4.6
16.4%
3
Image
Grok 4.7
12.1%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
Opus 5.5
64.8%
2
Image
Sonnet 5.5
61.8%
3
Image
GPT-6 Astra
58.2%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.