MBABench 2.0
MBABench 2.0 evaluates whether AI agents can build financial models in spreadsheets the way a finance professional would, across 101 tasks written from scratch by finance experts.
A task passes only if every numerical answer is right and the workbook meets a professional quality bar. MBABench is built by researchers at Columbia Business School and supported through Snorkel’s Open Benchmarks Grants.
At a glance
101
6
finance areas covered
16
financial model types
20
agents evaluated
Key takeaways
Astra in Codex led MBABench 2.0 at 15.8%, followed by Fable 5.1 in Claude Code and Astra in ChatGPT, both at 7.9%. Ten other agents scored between 1% and 8%, and nine agents scored 0%, including Sol in ChatGPT for Excel. Only Astra in Codex passed any of the 20 Hard tasks, at 5.0%.
The harness mattered as much as the model. Astra passed 15.8% of tasks in Codex but 7.9% in ChatGPT, and Fable 5.1 passed 7.9% in Claude Code but 2.0% in Claude for Excel. Getting the numbers right was the main obstacle: the most accurate agent, Fable 5.1 in Claude Code, got every answer right on 58.9% of tasks.
Leaderboard
| Rank | Agent | Harness | Effort | Pass rate | ± SE | Numerical accuracy | ± SE | Passed |
|---|---|---|---|---|---|---|---|---|
| 1 | Astra | Codex | xhigh |
15.8%
|
± 3.6 |
54.4%
|
± 5.3 | 16 / 101 |
| 2 | Fable 5.1 | Claude Code | max |
7.9%
|
± 2.7 |
58.9%
|
± 5.2 | 8 / 101 |
| 3 | Astra | ChatGPT | ultra |
7.9%
|
± 2.7 |
51.1%
|
± 5.3 | 8 / 101 |
| 4 | Fable 5.1 | Claude Cowork | max |
5%
|
± 2.2 |
56.7%
|
± 5.3 | 5 / 101 |
| 5 | Opus 5 | Claude Code | max |
4%
|
± 1.9 |
46.7%
|
± 5.3 | 4 / 101 |
| 6 | GPT-6-Pro | ChatGPT | n/a |
3%
|
± 1.7 |
48.9%
|
± 5.3 | 3 / 101 |
| 7 | Fable 5.1 | Claude for Excel | n/a |
2%
|
± 1.4 |
52.2%
|
± 5.3 | 2 / 101 |
| 8 | Opus 5 | Claude for Excel | n/a |
1%
|
± 1 |
54.4%
|
± 5.3 | 1 / 101 |
| 9 | Opus 5 | Claude Cowork | max |
1%
|
± 1 |
50%
|
± 5.3 | 1 / 101 |
| 10 | Grok 4.6 | Codex | xhigh |
1%
|
± 1 |
45.6%
|
± 5.3 | 1 / 101 |
| 11 | Qwen 3.8 | Codex | xhigh |
1%
|
± 1 |
40%
|
± 5.2 | 1 / 101 |
| 12 | Fable 5.1 | In-house harness | max |
0%
|
± 0 |
50%
|
± 5.3 | 0 / 101 |
| 13 | Sol | ChatGPT for Excel | xhigh |
0%
|
± 0 |
45.6%
|
± 5.3 | 0 / 101 |
| 14 | Gemini 3.8 | Codex | high |
0%
|
± 0 |
45.6%
|
± 5.3 | 0 / 101 |
| 15 | Kimi K3 | Codex | max |
0%
|
± 0 |
44.4%
|
± 5.3 | 0 / 101 |
| 16 | Astra | In-house harness | xhigh |
0%
|
± 0 |
41.1%
|
± 5.2 | 0 / 101 |
| 17 | Kimi K3 | In-house harness | max |
0%
|
± 0 |
36.7%
|
± 5.1 | 0 / 101 |
| 18 | Qwen 3.8 | In-house harness | xhigh |
0%
|
± 0 |
28.9%
|
± 4.8 | 0 / 101 |
| 19 | Gemini 3.8 | In-house harness | high |
0%
|
± 0 |
15.6%
|
± 3.8 | 0 / 101 |
| 20 | Grok 4.6 | In-house harness | xhigh |
0%
|
± 0 |
7.8%
|
± 2.8 | 0 / 101 |
| Rank | Agent | Harness | Effort | Pass rate | ± SE | Numerical accuracy | ± SE | Passed |
|---|---|---|---|---|---|---|---|---|
| 1 | Astra | ChatGPT | ultra |
54.5%
|
± 15 |
90.9%
|
± 8.7 | 6 / 11 |
| 2 | Fable 5.1 | Claude Code | max |
45.5%
|
± 15 |
81.8%
|
± 11.6 | 5 / 11 |
| 3 | Astra | Codex | xhigh |
45.5%
|
± 15 |
72.7%
|
± 13.4 | 5 / 11 |
| 4 | Opus 5 | Claude Code | max |
18.2%
|
± 11.6 |
72.7%
|
± 13.4 | 2 / 11 |
| 5 | GPT-6-Pro | ChatGPT | n/a |
18.2%
|
± 11.6 |
72.7%
|
± 13.4 | 2 / 11 |
| 6 | Fable 5.1 | Claude Cowork | max |
9.1%
|
± 8.7 |
90.9%
|
± 8.7 | 1 / 11 |
| 7 | Fable 5.1 | Claude for Excel | n/a |
9.1%
|
± 8.7 |
81.8%
|
± 11.6 | 1 / 11 |
| 8 | Opus 5 | Claude for Excel | n/a |
9.1%
|
± 8.7 |
72.7%
|
± 13.4 | 1 / 11 |
| 9 | Grok 4.6 | Codex | xhigh |
0%
|
± 0 |
90.9%
|
± 8.7 | 0 / 11 |
| 10 | Fable 5.1 | In-house harness | max |
0%
|
± 0 |
81.8%
|
± 11.6 | 0 / 11 |
| 11 | Opus 5 | Claude Cowork | max |
0%
|
± 0 |
72.7%
|
± 13.4 | 0 / 11 |
| 12 | Sol | ChatGPT for Excel | xhigh |
0%
|
± 0 |
72.7%
|
± 13.4 | 0 / 11 |
| 13 | Astra | In-house harness | xhigh |
0%
|
± 0 |
72.7%
|
± 13.4 | 0 / 11 |
| 14 | Qwen 3.8 | Codex | xhigh |
0%
|
± 0 |
63.6%
|
± 14.5 | 0 / 11 |
| 15 | Gemini 3.8 | Codex | high |
0%
|
± 0 |
63.6%
|
± 14.5 | 0 / 11 |
| 16 | Kimi K3 | In-house harness | max |
0%
|
± 0 |
54.5%
|
± 15 | 0 / 11 |
| 17 | Grok 4.6 | In-house harness | xhigh |
0%
|
± 0 |
54.5%
|
± 15 | 0 / 11 |
| 18 | Kimi K3 | Codex | max |
0%
|
± 0 |
45.5%
|
± 15 | 0 / 11 |
| 19 | Qwen 3.8 | In-house harness | xhigh |
0%
|
± 0 |
45.5%
|
± 15 | 0 / 11 |
| 20 | Gemini 3.8 | In-house harness | high |
0%
|
± 0 |
36.4%
|
± 14.5 | 0 / 11 |
| Rank | Agent | Harness | Effort | Pass rate | ± SE | Numerical accuracy | ± SE | Passed |
|---|---|---|---|---|---|---|---|---|
| 1 | Astra | Codex | xhigh |
22.5%
|
± 6.6 |
55%
|
± 7.9 | 9 / 40 |
| 2 | Fable 5.1 | Claude Cowork | max |
10%
|
± 4.7 |
60%
|
± 7.7 | 4 / 40 |
| 3 | Fable 5.1 | Claude Code | max |
5%
|
± 3.4 |
62.5%
|
± 7.7 | 2 / 40 |
| 4 | Astra | ChatGPT | ultra |
5%
|
± 3.4 |
52.5%
|
± 7.9 | 2 / 40 |
| 5 | Opus 5 | Claude Code | max |
5%
|
± 3.4 |
52.5%
|
± 7.9 | 2 / 40 |
| 6 | Grok 4.6 | Codex | xhigh |
2.5%
|
± 2.5 |
57.5%
|
± 7.8 | 1 / 40 |
| 7 | Fable 5.1 | Claude for Excel | n/a |
2.5%
|
± 2.5 |
55%
|
± 7.9 | 1 / 40 |
| 8 | GPT-6-Pro | ChatGPT | n/a |
2.5%
|
± 2.5 |
50%
|
± 7.9 | 1 / 40 |
| 9 | Qwen 3.8 | Codex | xhigh |
2.5%
|
± 2.5 |
42.5%
|
± 7.8 | 1 / 40 |
| 10 | Fable 5.1 | In-house harness | max |
0%
|
± 0 |
62.5%
|
± 7.7 | 0 / 40 |
| 11 | Opus 5 | Claude for Excel | n/a |
0%
|
± 0 |
57.5%
|
± 7.8 | 0 / 40 |
| 12 | Opus 5 | Claude Cowork | max |
0%
|
± 0 |
52.5%
|
± 7.9 | 0 / 40 |
| 13 | Kimi K3 | Codex | max |
0%
|
± 0 |
52.5%
|
± 7.9 | 0 / 40 |
| 14 | Astra | In-house harness | xhigh |
0%
|
± 0 |
52.5%
|
± 7.9 | 0 / 40 |
| 15 | Gemini 3.8 | Codex | high |
0%
|
± 0 |
52.5%
|
± 7.9 | 0 / 40 |
| 16 | Sol | ChatGPT for Excel | xhigh |
0%
|
± 0 |
47.5%
|
± 7.9 | 0 / 40 |
| 17 | Kimi K3 | In-house harness | max |
0%
|
± 0 |
45%
|
± 7.9 | 0 / 40 |
| 18 | Qwen 3.8 | In-house harness | xhigh |
0%
|
± 0 |
40%
|
± 7.7 | 0 / 40 |
| 19 | Gemini 3.8 | In-house harness | high |
0%
|
± 0 |
22.5%
|
± 6.6 | 0 / 40 |
| 20 | Grok 4.6 | In-house harness | xhigh |
0%
|
± 0 |
10%
|
± 4.7 | 0 / 40 |
| Rank | Agent | Harness | Effort | Pass rate | ± SE | Numerical accuracy | ± SE | Passed |
|---|---|---|---|---|---|---|---|---|
| 1 | Fable 5.1 | Claude Code | max |
3.3%
|
± 3.3 |
63.3%
|
± 8.8 | 1 / 30 |
| 2 | Astra | Codex | xhigh |
3.3%
|
± 3.3 |
60%
|
± 8.9 | 1 / 30 |
| 3 | Opus 5 | Claude Cowork | max |
3.3%
|
± 3.3 |
53.3%
|
± 9.1 | 1 / 30 |
| 4 | Fable 5.1 | Claude Cowork | max |
0%
|
± 0 |
63.3%
|
± 8.8 | 0 / 30 |
| 5 | Opus 5 | Claude for Excel | n/a |
0%
|
± 0 |
60%
|
± 8.9 | 0 / 30 |
| 6 | Astra | ChatGPT | ultra |
0%
|
± 0 |
56.7%
|
± 9 | 0 / 30 |
| 7 | GPT-6-Pro | ChatGPT | n/a |
0%
|
± 0 |
56.7%
|
± 9 | 0 / 30 |
| 8 | Fable 5.1 | Claude for Excel | n/a |
0%
|
± 0 |
56.7%
|
± 9 | 0 / 30 |
| 9 | Opus 5 | Claude Code | max |
0%
|
± 0 |
53.3%
|
± 9.1 | 0 / 30 |
| 10 | Sol | ChatGPT for Excel | xhigh |
0%
|
± 0 |
53.3%
|
± 9.1 | 0 / 30 |
| 11 | Fable 5.1 | In-house harness | max |
0%
|
± 0 |
50%
|
± 9.1 | 0 / 30 |
| 12 | Astra | In-house harness | xhigh |
0%
|
± 0 |
50%
|
± 9.1 | 0 / 30 |
| 13 | Gemini 3.8 | Codex | high |
0%
|
± 0 |
50%
|
± 9.1 | 0 / 30 |
| 14 | Grok 4.6 | Codex | xhigh |
0%
|
± 0 |
46.7%
|
± 9.1 | 0 / 30 |
| 15 | Kimi K3 | In-house harness | max |
0%
|
± 0 |
43.3%
|
± 9 | 0 / 30 |
| 16 | Qwen 3.8 | Codex | xhigh |
0%
|
± 0 |
40%
|
± 8.9 | 0 / 30 |
| 17 | Kimi K3 | Codex | max |
0%
|
± 0 |
40%
|
± 8.9 | 0 / 30 |
| 18 | Qwen 3.8 | In-house harness | xhigh |
0%
|
± 0 |
30%
|
± 8.4 | 0 / 30 |
| 19 | Gemini 3.8 | In-house harness | high |
0%
|
± 0 |
16.7%
|
± 6.8 | 0 / 30 |
| 20 | Grok 4.6 | In-house harness | xhigh |
0%
|
± 0 |
10%
|
± 5.5 | 0 / 30 |
| Rank | Agent | Harness | Effort | Pass rate | ± SE | Numerical accuracy | ± SE | Passed |
|---|---|---|---|---|---|---|---|---|
| 1 | Astra | Codex | xhigh |
5%
|
± 4.9 |
45%
|
± 11.1 | 1 / 20 |
| 2 | Fable 5.1 | Claude Code | max |
0%
|
± 0 |
45%
|
± 11.1 | 0 / 20 |
| 3 | Astra | ChatGPT | ultra |
0%
|
± 0 |
40%
|
± 11 | 0 / 20 |
| 4 | Fable 5.1 | Claude Cowork | max |
0%
|
± 0 |
40%
|
± 11 | 0 / 20 |
| 5 | Fable 5.1 | Claude for Excel | n/a |
0%
|
± 0 |
40%
|
± 11 | 0 / 20 |
| 6 | Opus 5 | Claude Cowork | max |
0%
|
± 0 |
40%
|
± 11 | 0 / 20 |
| 7 | Opus 5 | Claude for Excel | n/a |
0%
|
± 0 |
40%
|
± 11 | 0 / 20 |
| 8 | GPT-6-Pro | ChatGPT | n/a |
0%
|
± 0 |
35%
|
± 10.7 | 0 / 20 |
| 9 | Qwen 3.8 | Codex | xhigh |
0%
|
± 0 |
35%
|
± 10.7 | 0 / 20 |
| 10 | Kimi K3 | Codex | max |
0%
|
± 0 |
35%
|
± 10.7 | 0 / 20 |
| 11 | Sol | ChatGPT for Excel | xhigh |
0%
|
± 0 |
30%
|
± 10.2 | 0 / 20 |
| 12 | Opus 5 | Claude Code | max |
0%
|
± 0 |
25%
|
± 9.7 | 0 / 20 |
| 13 | Fable 5.1 | In-house harness | max |
0%
|
± 0 |
25%
|
± 9.7 | 0 / 20 |
| 14 | Gemini 3.8 | Codex | high |
0%
|
± 0 |
25%
|
± 9.7 | 0 / 20 |
| 15 | Grok 4.6 | Codex | xhigh |
0%
|
± 0 |
20%
|
± 8.9 | 0 / 20 |
| 16 | Kimi K3 | In-house harness | max |
0%
|
± 0 |
10%
|
± 6.7 | 0 / 20 |
| 17 | Astra | In-house harness | xhigh |
0%
|
± 0 |
5%
|
± 4.9 | 0 / 20 |
| 18 | Qwen 3.8 | In-house harness | xhigh |
0%
|
± 0 |
5%
|
± 4.9 | 0 / 20 |
| 19 | Grok 4.6 | In-house harness | xhigh |
0%
|
± 0 |
0%
|
± 0 | 0 / 20 |
| 20 | Gemini 3.8 | In-house harness | high |
0%
|
± 0 |
0%
|
± 0 | 0 / 20 |
Pass: every numerical answer is right and the workbook has at most 1, 2, 3 and 2 other criteria flagged across the four importance tiers. Accuracy: every answer matches the key; Overall accuracy is measured on the 90 tasks rated Medium or harder. ± is one standard error.
Methodology
Evaluator
Each agent is run in its own harness, such as Codex, Claude Code, ChatGPT or Claude for Excel. The same model can score very differently depending on the harness, so the leaderboard lists each pairing separately.
scoring note
Pass rate is strict. A task counts only when every numerical answer is right and the workbook meets the quality bar, so an agent can be mostly accurate and still score 0%.
Behind the benchmark
Financial modeling is more than arithmetic. A correct model needs the right structure, linked formulas and assumptions that hold together across the workbook, so one wrong cell can change every answer downstream.
MBABench 2.0 was built by researchers at Columbia Business School, led by Hongseok Namkoong and Thomson Yen. Its 101 tasks were written from scratch by finance experts and span 6 finance areas and 16 financial model types.
MBABench is supported through Snorkel’s Open Benchmarks Grants. Tasks, data and code are available on the MBABench website and GitHub.

