Open Benchmarks Grants
agent-le-logo

Agents' Last Exam

A benchmark for evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes. 55 sub-industries, 1,500+ tasks toward a 5,000-task target, sourced and validated by 300+ industry experts.

Built with
ImageImageSnorkel AI logo lockup mono white outline png
Overview

Agents’ Last Exam (ALE) is building the broadest-coverage agent evaluation benchmark to date, measuring performance on long-horizon, economically valuable tasks with verifiable outcomes. The benchmark covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy), spanning all 55 targeted sub-industries.

ALE-V1 ships 147 reference tasks across 55 industries as the current public subset of a 1,500+ task corpus. Many tasks require private data or licensed software and remain in a separate private pool. ALE uses rolling evaluation: every ~6 months a new public subset is published with fresh instances, while private tasks rotate in and retired public tasks rotate out, to limit benchmark leakage.

Leaderboard

Rank Model Effort Harness Pass Rate Score Est. Cost Runtime Input Tokens Output Tokens
1 GPT-6 Astra Codex
34.2%
59.3%
$948 146h 16m 510.3M 5.8M
2 Muse Spark 1.3 Codex
32.2%
55.8%
$428 80h 37m 2.1B 11.8M
3 Claude Opus 5 Claude Code
31.6%
55.2%
$1,638 128h 29m 1.8B 10.1M
4 GPT-5.6 Sol Codex
30.6%
53.6%
$772 94h 39m 762.9M 3.8M
5 GPT-5.6 Luna Codex
30.3%
49.4%
$235 66h 7m 1.4B 4.2M
6 Kimi K3 Max Kimi Code
28.3%
51.6%
$606 186h 5m 1.4B 7.5M
7 GPT-5.6 Terra Codex
28%
50.7%
$545 118h 39m 1.2B 5.8M
8 Qwen3 8-Max XHigh Claude Code
27%
52.5%
$486 202h 41m 1.2B 12.4M
9 Grok 4.5 High Grok Build
27%
51.2%
$337 89h 8m 668.3M 5.4M
10 Kimi K3 Max Claude Code
27%
50.7%
$464 215h 32m 950.6M 7.4M
11 Claude Opus 4.8 Claude Code
27%
45.1%
$3,985 179h 18m 1.7B 22.1M
12 GPT-5.5 Codex
26.6%
47.9%
$602 97h 8m 560.5M 6.4M
13 Claude Fable 5 Claude Code
25.7%
48.7%
$4,340 70h 0m 2.7B 10.4M
14 GPT-5.5 High ALE-Claw
23%
45.8%
$310 35h 34m 334.3M 2.3M
15 GPT-5.5 High OpenClaw
21.7%
42.1%
$447 82h 38m 469.3M 3.3M
16 GPT-5.5 Medium Cursor CLI
21.4%
40.7%
$174 58h 47m 148.1M 1.7M
17 Claude Opus 4.7 High Cursor CLI
21.1%
42.5%
$1,899 63h 49m 429.5M 4.5M
18 GPT-5.4 High OpenClaw
20.5%
37.3%
$274 127h 46m 488.8M 7.3M
19 Qwen3 8-27B Max Claude Code
20.4%
42.9%
277h 59m 2.7B 15.0M
20 GLM-5.2 Max Claude Code
20.4%
41.1%
$1,086 107h 42m 1.3B 7.7M
21 Composer 2.5 Adaptive Cursor CLI
20.4%
39.5%
$177 77h 48m 338.8M 2.9M
22 GPT-5.5 High Droid
19.7%
39.2%
$242 63h 49m 234.8M 2.2M
23 Seed 2.1 Pro - Claude Code
19.5%
41.9%
$936 155h 14m 2.1B 8.4M
24 Claude Opus 4.7 High ALE-Claw
18.4%
40.5%
$1,132 77h 39m 1.3B 5.6M
25 Gemini 3.1 Pro High Gemini CLI
16.4%
32.7%
$2,018 92h 21m 1.2B 3.5M
26 Claude Opus 4.7 High OpenClaw
15.8%
35.2%
$1,689 124h 10m 807.3M 4.0M
27 Gemini 3.1 Pro High OpenClaw
14.1%
29.3%
$2,969 147h 6m 3.4B 3.7M
28 Claude Opus 4.7 High Claude Code
13.8%
35.8%
$1,793 42h 36m 456.4M 3.7M
29 Claude Opus 4.7 High Droid
13.5%
31.7%
$1,356 28h 16m 352.1M 2.7M
30 Seed 2.1 Pro High OpenClaw
13.1%
32.7%
$440 190h 18m 1.6B 20.4M
31 DeepSeek V4 Pro High OpenClaw
12.4%
27.6%
$273 130h 47m 511.7M 4.7M
32 Qwen3 7-Max High OpenClaw
11.8%
31.6%
$664 101h 0m 1.4B 17.6M
33 GPT-5.4 High ALE-Claw
11.8%
28.6%
$335 55h 28m 1.1B 2.0M
34 GLM-5.1 High OpenClaw
11.5%
28.1%
$440 168h 39m 755.8M 5.9M
35 Kimi K2.6 High OpenClaw
9.2%
21.7%
$123 183h 18m 295.9M 6.0M
36 Qwen3 6-Plus High OpenClaw
8.6%
24.6%
$254 134h 43m 738.9M 7.0M
37 MiMo V2.5 High OpenClaw
8.6%
23.6%
$43 107h 22m 394.9M 4.0M
38 Grok 4.3 - Grok CLI
7.2%
21.3%
$285 43h 7m 223.1M 2.3M
39 MiniMax M2.7 High OpenClaw
5.9%
14.2%
$27 144h 37m 277.2M 4.4M
40 Grok 4.3 High OpenClaw
4.3%
15.8%
$80 79h 4m 168.7M 2.7M
Rank Model Effort Harness Pass Rate Score Est. Cost Runtime Input Tokens Output Tokens
1 GPT-6 Astra Codex
33.3%
61.4%
$507 70h 32m 227.5M 3.7M
2 Muse Spark 1.3 Codex
33.3%
57.5%
$256 37h 32m 1.3B 7.7M
3 Claude Opus 5 Claude Code
29.5%
57.3%
$857 73h 44m 819.5M 6.1M
4 GPT-5.6 Luna Codex
29.5%
50.7%
$144 24h 53m 910.7M 2.8M
5 GPT-5.6 Sol Codex
28.6%
53.6%
$446 49h 46m 443.5M 2.5M
6 Kimi K3 Max Claude Code
27.6%
52.8%
$251 109h 24m 465.2M 4.7M
7 Kimi K3 Max Kimi Code
27.6%
52%
$305 102h 15m 604.4M 5.0M
8 GPT-5.5 Codex
27.1%
49.2%
$385 47h 34m 357.1M 4.3M
9 GPT-5.6 Terra Codex
26.2%
49.4%
$235 49h 6m 528.0M 2.6M
10 Grok 4.5 High Grok Build
25.7%
50.6%
$139 51h 49m 247.6M 3.3M
11 Claude Opus 4.8 Claude Code
25.7%
44.1%
$1,953 110h 57m 642.2M 15.3M
12 Qwen3 8-Max XHigh Claude Code
24.8%
51.1%
$249 113h 20m 559.9M 6.8M
13 GPT-5.5 High ALE-Claw
24.8%
48.1%
$109 16h 44m 77.0M 1.3M
14 GPT-5.5 High OpenClaw
24.8%
46.1%
$219 31h 30m 211.2M 2.2M
15 Claude Fable 5 Claude Code
23.8%
48.5%
$1,663 34h 35m 900.2M 6.4M
16 GLM-5.2 Max Claude Code
23.8%
44.2%
$981 64h 18m 1.2B 6.3M
17 Claude Opus 4.7 High Cursor CLI
22.9%
45.7%
$1,045 24h 54m 207.5M 2.5M
18 GPT-5.5 Medium Cursor CLI
22.9%
43.4%
$97 14h 30m 68.5M 1.1M
19 GPT-5.4 High OpenClaw
22.5%
41.8%
$169 50h 7m 285.8M 4.9M
20 Claude Opus 4.7 High ALE-Claw
20%
43.3%
$416 44h 15m 389.6M 2.9M
21 GPT-5.5 High Droid
20%
40.5%
$136 37h 27m 117.0M 1.5M
22 Qwen3 8-27B Max Claude Code
19%
43.5%
162h 13m 1.6B 10.2M
23 Composer 2.5 Adaptive Cursor CLI
19%
42.2%
$81 60h 10m 153.2M 1.8M
24 Seed 2.1 Pro - Claude Code
19%
42.2%
$602 88h 11m 1.1B 5.8M
25 Gemini 3.1 Pro High Gemini CLI
18.1%
37.2%
$1,010 52h 27m 450.8M 2.4M
26 Claude Opus 4.7 High OpenClaw
17.1%
38.7%
$960 53h 29m 516.2M 2.8M
27 Gemini 3.1 Pro High OpenClaw
15.7%
32.4%
$1,303 66h 23m 1.4B 2.3M
28 Claude Opus 4.7 High Claude Code
15.2%
38.9%
$1,165 25h 47m 281.0M 2.7M
29 Claude Opus 4.7 High Droid
15.2%
34.7%
$710 19h 17m 176.9M 1.7M
30 DeepSeek V4 Pro High OpenClaw
14.1%
30.8%
$161 56h 12m 312.8M 3.3M
31 Qwen3 7-Max High OpenClaw
13.3%
34.6%
$534 34h 23m 1.0B 15.8M
32 GPT-5.4 High ALE-Claw
13.3%
33.8%
$168 26h 31m 541.3M 1.1M
33 Claude Sonnet 4.6 Medium Forgecode
13.3%
29%
$100 51h 31m 136.5M 1.7M
34 Claude Sonnet 4.6 Medium Hermes
13.2%
32%
$437 25h 12m 195.6M 2.0M
35 GLM-5.1 High OpenClaw
12.9%
30.8%
$213 73h 57m 361.0M 3.3M
36 Claude Sonnet 4.6 Off Terminus 2
11.9%
30.9%
$320 73h 1m 760.4M 3.1M
37 Seed 2.1 Pro High OpenClaw
11.4%
33.9%
$440 98h 0m 1.4B 18.4M
38 Claude Sonnet 4.6 High OpenClaw
11.4%
31.6%
$181 33h 51m 247.0M 1.9M
39 Qwen3 6-Plus High OpenClaw
10.5%
29%
$128 62h 28m 366.8M 4.4M
40 MiMo V2.5 High OpenClaw
10%
26.5%
$32 51h 24m 286.4M 3.0M
41 Claude Sonnet 4.6 Not reported Openhands
9%
19.8%
$243 35h 47m 349.9M 4.4M
42 Grok 4.3 - Grok CLI
8.6%
25.9%
$205 36h 12m 161.0M 1.7M
43 Kimi K2.6 High OpenClaw
8.1%
21.2%
$91 89h 31m 222.9M 4.4M
44 MiniMax M2.7 High OpenClaw
5.7%
14.6%
$22 52h 56m 242.7M 3.2M
45 Grok 4.3 High OpenClaw
4.3%
17.8%
$61 36h 8m 135.8M 2.1M
Rank Model Effort Harness Pass Rate Score Est. Cost Runtime Input Tokens Output Tokens
1 GPT-6 Astra Codex
52.2%
82.3%
$272 25h 51m 150.6M 1.4M
2 GPT-5.6 Luna Codex
49.3%
75.6%
$69 19h 22m 359.7M 1.6M
3 GPT-5.6 Sol Codex
47.8%
78.8%
$322 40h 29m 304.2M 1.8M
4 Claude Opus 5 Claude Code
47.8%
78.6%
$430 30h 25m 470.8M 3.2M
5 Muse Spark 1.3 Codex
47.8%
77.8%
$103 18h 20m 458.0M 3.8M
6 GPT-5.6 Terra Codex
44.8%
73.9%
$96 31h 27m 170.7M 1.3M
7 GPT-5.5 Codex
44%
72.4%
$170 29h 32m 126.2M 2.3M
8 Claude Opus 4.8 Claude Code
43.3%
64%
$1,106 49h 40m 310.5M 8.3M
9 Kimi K3 Max Claude Code
40.3%
71.6%
$145 66h 36m 270.2M 2.7M
10 Qwen3 8-Max XHigh Claude Code
38.8%
72.1%
$137 62h 29m 322.0M 4.7M
11 Grok 4.5 High Grok Build
38.8%
71.9%
$121 20h 54m 239.6M 2.0M
12 Kimi K3 Max Kimi Code
38.8%
69.7%
$129 50h 48m 256.7M 2.2M
13 Claude Fable 5 Claude Code
37.3%
71.1%
$1,018 21h 14m 504.9M 3.6M
14 GPT-5.5 High OpenClaw
37.3%
67.2%
$218 41h 19m 223.0M 1.5M
15 Composer 2.5 Adaptive Cursor CLI
34.3%
62.3%
$50 22h 1m 94.1M 1.1M
16 GPT-5.5 Medium Cursor CLI
33.6%
62.3%
$68 30h 10m 51.8M 741.1K
17 GPT-5.4 High OpenClaw
33.6%
57.8%
$69 40h 56m 86.0M 2.4M
18 GPT-5.5 High ALE-Claw
32.8%
67.4%
$148 14h 57m 167.2M 1.0M
19 GLM-5.2 Max Claude Code
32.8%
60.3%
$204 32h 45m 187.0M 2.7M
20 Claude Opus 4.7 High Cursor CLI
31.3%
62.7%
$558 33h 26m 110.3M 1.6M
21 Qwen3 8-27B Max Claude Code
31.3%
60%
82h 54m 684.6M 6.1M
22 GPT-5.5 High Droid
29.9%
58.2%
$106 23h 26m 98.4M 962.6K
23 Claude Opus 4.7 High Droid
29.1%
61.7%
$738 16h 6m 176.2M 1.5M
24 Claude Opus 4.7 High ALE-Claw
28.4%
60.5%
$312 20h 40m 359.2M 1.9M
25 Claude Opus 4.7 High OpenClaw
28.4%
58%
$508 47h 42m 201.7M 1.5M
26 Gemini 3.1 Pro High Gemini CLI
28.4%
54.9%
$342 27h 40m 239.4M 1.1M
27 Seed 2.1 Pro - Claude Code
26.9%
60.1%
$376 66h 30m 697.7M 3.1M
28 Gemini 3.1 Pro High OpenClaw
26.1%
49.5%
$575 48h 37m 616.7M 938.9K
29 Claude Opus 4.7 High Claude Code
22.4%
55.8%
$496 16h 17m 120.4M 1.1M
30 GPT-5.4 High ALE-Claw
20.9%
45.4%
$66 17h 22m 178.6M 684.8K
31 GLM-5.1 High OpenClaw
20.1%
45.6%
$108 62h 17m 183.5M 1.6M
32 DeepSeek V4 Pro High OpenClaw
19.9%
43.8%
$109 58h 50m 208.2M 1.9M
33 Qwen3 7-Max High OpenClaw
17.9%
46.9%
$247 38h 11m 502.0M 7.3M
34 Seed 2.1 Pro High OpenClaw
17.9%
46.9%
$201 76h 7m 640.9M 8.9M
35 Kimi K2.6 High OpenClaw
15.7%
35.1%
$47 61h 6m 118.5M 2.1M
36 Qwen3 6-Plus High OpenClaw
12.7%
36.4%
$90 60h 13m 259.7M 2.7M
37 MiMo V2.5 High OpenClaw
11.9%
35.1%
$13 43h 37m 105.4M 1.5M
38 Grok 4.3 - Grok CLI
10.4%
31.8%
$127 21h 59m 99.8M 1.0M
39 MiniMax M2.7 High OpenClaw
10.4%
24.5%
$9 63h 22m 98.2M 1.4M
40 Grok 4.3 High OpenClaw
6.7%
24.4%
$35 31h 15m 73.8M 1.2M
Rank Model Effort Harness Pass Rate Score Est. Cost Runtime Input Tokens Output Tokens
1 GPT-6 Astra Codex
30.9%
51.8%
$252 44h 47m 122.0M 1.7M
2 GPT-5.6 Sol Codex
30%
46.6%
$256 28h 47m 235.4M 1.3M
3 Kimi K3 Max Kimi Code
29.1%
49.9%
$206 55h 44m 465.4M 2.8M
4 Claude Opus 5 Claude Code
29.1%
47.9%
$527 39h 3m 586.2M 3.3M
5 Muse Spark 1.3 Codex
29.1%
47.4%
$130 18h 18m 623.1M 3.8M
6 GPT-5.6 Luna Codex
29.1%
42.5%
$107 19h 52m 652.4M 2.2M
7 GPT-5.6 Terra Codex
27.3%
46.2%
$150 38h 15m 316.9M 1.8M
8 Claude Fable 5 Claude Code
25.5%
42%
$1,826 25h 34m 1.1B 3.8M
9 Seed 2.1 Pro - Claude Code
24.1%
40.1%
$289 42h 44m 639.5M 3.2M
10 Qwen3 8-Max XHigh Claude Code
23.6%
46.4%
$149 57h 33m 392.1M 3.6M
11 Kimi K3 Max Claude Code
23.6%
44.9%
$144 64h 56m 297.8M 2.4M
12 Grok 4.5 High Grok Build
23.6%
44.6%
$114 29h 56m 230.1M 2.1M
13 Claude Opus 4.8 Claude Code
23.6%
43.1%
$1,324 59h 37m 524.3M 8.2M
14 GPT-5.5 High ALE-Claw
23.6%
41.1%
$70 8h 28m 54.9M 734.8K
15 GPT-5.5 Codex
22.4%
37.6%
$130 20h 57m 130.8M 1.2M
16 Qwen3 8-27B Max Claude Code
21.8%
42.5%
98h 57m 1.2B 5.4M
17 Claude Opus 4.7 High Cursor CLI
20%
37.6%
$842 14h 48m 215.6M 1.9M
18 GPT-5.5 Medium Cursor CLI
20%
34%
$50 17h 24m 36.4M 552.1K
19 GPT-5.4 High OpenClaw
19.4%
34.3%
$123 57h 48m 239.6M 2.9M
20 GLM-5.2 Max Claude Code
18.2%
37%
$390 35h 45m 338.8M 2.7M
21 Claude Opus 4.7 High ALE-Claw
18.2%
36.6%
$246 15h 31m 288.8M 1.7M
22 GPT-5.5 High Droid
18.2%
35.2%
$69 15h 37m 58.6M 792.2K
23 GPT-5.5 High OpenClaw
18.2%
33.4%
$101 22h 58m 98.6M 1.1M
24 Composer 2.5 Adaptive Cursor CLI
18.2%
32%
$68 25h 41m 130.0M 1.1M
25 Seed 2.1 Pro High OpenClaw
13.5%
24.9%
$171 46h 53m 572.9M 7.7M
26 Claude Opus 4.7 High Claude Code
12.7%
29.1%
$747 12h 5m 202.0M 1.5M
27 Gemini 3.1 Pro High Gemini CLI
12.7%
26.4%
$962 25h 34m 481.7M 1.7M
28 Qwen3 7-Max High OpenClaw
10.9%
28.5%
$280 35h 46m 585.1M 7.1M
29 Claude Opus 4.7 High OpenClaw
10.9%
27.5%
$370 29h 47m 163.0M 1.1M
30 DeepSeek V4 Pro High OpenClaw
10.9%
23.8%
$68 38h 12m 114.5M 1.5M
31 Gemini 3.1 Pro High OpenClaw
10.9%
23.6%
$1,392 62h 47m 1.6B 1.7M
32 GPT-5.4 High ALE-Claw
9.1%
22.9%
$179 22h 36m 595.7M 930.0K
33 GLM-5.1 High OpenClaw
9.1%
21.7%
$165 61h 57m 282.4M 2.2M
34 MiMo V2.5 High OpenClaw
9.1%
20.8%
$17 34h 51m 159.9M 1.4M
35 Qwen3 6-Plus High OpenClaw
8.2%
22.9%
$111 46h 37m 322.1M 3.1M
36 Grok 4.3 - Grok CLI
7.3%
18.3%
$95 10h 30m 74.0M 896.2K
37 Kimi K2.6 High OpenClaw
6.4%
18.2%
$41 73h 35m 93.6M 2.4M
38 Grok 4.3 High OpenClaw
3.6%
13.5%
$25 24h 45m 51.6M 1.0M
39 Claude Opus 4.7 High Droid
3.6%
10.9%
$167 3h 8m 39.9M 472.1K
40 MiniMax M2.7 High OpenClaw
3.6%
8.4%
$10 46h 17m 95.2M 1.7M
Rank Model Effort Harness Pass Rate Score Est. Cost Runtime Input Tokens Output Tokens
1 Claude Opus 5 Claude Code
13.2%
19.7%
$952 75h 5m 1.0B 5.1M
2 GPT-6 Astra Codex
10.5%
27.4%
$310 39h 33m 203.6M 1.3M
3 Qwen3 8-Max XHigh Claude Code
10.5%
24.1%
$217 95h 41m 559.3M 4.8M
4 Muse Spark 1.3 Codex
10.5%
23.8%
$225 39h 4m 1.2B 4.9M
5 Grok 4.5 High Grok Build
10.5%
23.7%
$115 47h 43m 223.7M 1.7M
6 Kimi K3 Max Kimi Code
10.5%
20.6%
$308 91h 53m 748.2M 3.1M
7 Kimi K3 Max Claude Code
7.9%
20.5%
$200 100h 52m 437.6M 2.6M
8 Claude Fable 5 Claude Code
7.9%
16.5%
$2,091 30h 16m 1.6B 3.8M
9 GPT-5.6 Sol Codex
5.3%
19.4%
$325 38h 53m 352.3M 1.4M
10 GPT-5.6 Luna Codex
2.6%
15%
$197 31h 35m 1.4B 2.6M
11 GPT-5.6 Terra Codex
2.6%
14.9%
$264 50h 55m 678.1M 2.2M
12 Qwen3 8-27B Max Claude Code
2.6%
14.7%
114h 9m 894.8M 4.1M
13 Claude Opus 4.8 Claude Code
2.6%
14.4%
$1,805 84h 44m 961.0M 7.3M
14 GPT-5.5 Codex
2.6%
14.2%
$167 41h 36m 153.9M 905.0K
15 GLM-5.2 Max Claude Code
2.6%
12.9%
$638 47h 56m 782.7M 2.7M
16 GPT-5.5 High ALE-Claw
2.6%
12.8%
$107 13h 44m 126.3M 704.4K
17 Claude Opus 4.7 High Cursor CLI
2.6%
11.4%
$588 18h 31m 121.0M 1.2M
18 GPT-5.5 High Droid
2.6%
11.3%
$76 26h 27m 83.6M 551.8K
19 GPT-5.5 Medium Cursor CLI
2.6%
10.7%
$61 17h 12m 63.9M 456.4K
20 GPT-5.5 High OpenClaw
0%
10.9%
$155 24h 52m 178.2M 997.0K
21 Seed 2.1 Pro - Claude Code
0%
10.8%
$319 52h 3m 814.8M 2.8M
22 Seed 2.1 Pro High OpenClaw
0%
10.4%
$85 75h 30m 455.8M 4.8M
23 Claude Opus 4.7 High Claude Code
0%
10%
$685 17h 6m 170.1M 1.4M
24 Composer 2.5 Adaptive Cursor CLI
0%
8.8%
$78 34h 44m 151.0M 948.3K
25 GPT-5.4 High ALE-Claw
0%
8.1%
$100 17h 39m 321.4M 540.7K
26 Claude Opus 4.7 High Droid
0%
8%
$496 10h 18m 145.7M 920.1K
27 Claude Opus 4.7 High ALE-Claw
0%
7.9%
$594 43h 8m 707.3M 2.2M
28 GPT-5.4 High OpenClaw
0%
7.3%
$107 36h 16m 206.9M 2.5M
29 Qwen3 7-Max High OpenClaw
0%
6.4%
$173 33h 3m 377.8M 4.1M
30 GLM-5.1 High OpenClaw
0%
6.2%
$204 57h 10m 354.0M 2.5M
31 Qwen3 6-Plus High OpenClaw
0%
5%
$78 35h 46m 228.9M 1.9M
32 Claude Opus 4.7 High OpenClaw
0%
4.3%
$842 50h 0m 452.9M 1.5M
33 MiMo V2.5 High OpenClaw
0%
3.4%
$15 33h 3m 135.8M 1.3M
34 Gemini 3.1 Pro High OpenClaw
0%
3.1%
$1,200 45h 9m 1.4B 1.1M
35 DeepSeek V4 Pro High OpenClaw
0%
2.5%
$108 39h 7m 208.8M 1.7M
36 Grok 4.3 - Grok CLI
0%
2.3%
$71 11h 35m 56.2M 477.3K
37 Grok 4.3 High OpenClaw
0%
2.3%
$22 25h 4m 47.2M 621.3K
38 Kimi K2.6 High OpenClaw
0%
1.6%
$39 60h 46m 88.1M 1.8M
39 MiniMax M2.7 High OpenClaw
0%
1.3%
$9 44h 14m 87.7M 1.5M
40 Gemini 3.1 Pro High Gemini CLI
0%
0.9%
$733 40h 2m 497.3M 821.0K

*claude-fable-5: the variant Anthropic served during evaluation may differ from the published model's full capability tier, and re-runs cannot guarantee the higher-tier variant is selected. These numbers may understate the model's true ceiling. Learn more.

Pass rate vs input tokens

Efficiency frontier for this evaluation split. Up and to the left is better (higher pass rate at lower input-token spend). Several harness/model pairs reach the top of the split at a fraction of the cost of others; some pairs spend an order of magnitude more tokens without a corresponding pass-rate gain.

Hover any dot for the harness, model, and exact metrics. 

Industry coverage

Six representative task families across the 55 sub-industries ALE covers.
Image

Motion & VFX

Animation and visual effects production tasks in Adobe After Effects.
Image

3D modeling

3D model creation and editing tasks in Siemens NX.
Image

Game development

Scene setup, asset placement, and rendering tasks in Unreal Engine.
Image

Mold flow analysis

Simulation and mold flow analysis tasks in Moldex3D manufacturing software.
Image

Architectural modeling

3D modeling and energy analysis workflows in Rhino 3D for urban design.
Image

Brain imaging

Neuroimaging analysis and brain structure segmentation tasks in FSLeyes.

Sample tasks

A selection from the 147 public ALE-V1 tasks across 14 task categories. Each task ships with a sandboxed environment, a hidden reference, and a deterministic grader. Slugs link to the task source.

business finance

sec_10k_financial_parsing

Parse a SEC 10-K filing into a structured financial schema. Multi-step extraction, table normalization, and cross-reference validation against the original document.

business finance

financial_stmt_reconstruction_aapl_fy2024
Reconstruct Apple’s FY2024 financial statement from primary disclosure documents. Validates whether the agent surfaces the exact reported figures and footnote-relevant adjustments.
engineering
mold-flow / 220089
Set up a Moldex3D mold-flow simulation, run it to convergence, and report fill time / pressure metrics matching the held-out reference run.

health medicine

Clinical_Variant_Annotation

Annotate a clinical variant set using standard pipelines (VEP, ClinVar, etc.) and produce a report graded against a curated reference.

life sciences

WGS_Variant_Calling
Run a whole-genome sequencing variant-calling pipeline and produce VCF output. Scored on precision and recall against a held-out truth VCF.

computing math

k8s_payment_api_root_cause_analysis
Diagnose a failing payment API in a Kubernetes cluster. Multi-hop investigation across logs, metrics, manifests, and traces, scored on the correct root-cause identification.

visual media

video_storyboard_001
Build a shot-by-shot video storyboard from a brief, formatted to industry conventions. Graded on coverage, continuity, and adherence to the reference shot list.
legal
legal_dr_fees_01
Compute legal fees from a billing register according to jurisdictional rules. Tests structured extraction plus rule-following against an authoritative reference total.

Methodology

Metrics

Pass Rate — fraction of tasks the agent fully completed (strict success). Score — average graded outcome across all tasks, including partial credit. Both computed by deterministic graders against hidden references.

Verifiable Outcomes
Hidden references plus deterministic graders, not LLM-as-a-judge. Tasks sourced from real professional workflows (After Effects, Siemens NX, Unreal Engine, Moldex3D, Rhino 3D, FSLeyes, and 49 more applications) and validated by domain experts before inclusion.
Rolling Evaluation

Every ~6 months, a new public subset releases with fresh instances. Private tasks rotate into the public pool, retired public tasks rotate out, and held-out private tasks score the official leaderboard, to limit benchmark leakage.

Reference Harnesses

Two open harnesses ship with the framework: the official Claude Code CLI and the in-tree OpenClaw harness. Submissions also include Codex, Cursor CLI, Droid, Gemini CLI, Grok CLI, and the ALE Claw reference harness.

Acknowledgments

Agents’ Last Exam is co-led by UC Berkeley RDI and the RDI Foundation, with funding support and contributions from Snorkel AI via the Open Benchmarks Grants program. The benchmark draws task contributions from 300+ industry experts across 44 academic institutions (MIT, Harvard, Stanford, UC Berkeley, Oxford, CMU, Caltech, ETH Zurich, Yale, Columbia, and more) and industry organizations including Goldman Sachs, JPMorgan, Morgan Stanley, PIMCO, Meta, Amazon, Adobe, Oracle, Hippocratic AI, and HubSpot.

Advisory Committee includes George Em Karniadakis (Brown), Tapio Schneider (Caltech), Teresa Head-Gordon (UC Berkeley), Laure Zanna (NYU), Jack Gallant (UC Berkeley), Tarek Zohdi (UC Berkeley), Ida Sim (UCSF), Arvind Rao (U Michigan), Kaan Ozbay (NYU), Carl Boettiger (UC Berkeley), Kyle Steinfeld (UC Berkeley), Yamini Rangan (HubSpot), and Bradley Rothenberg (nTop).

FAQs

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

Open Benchmarks Grants

Terminal-Bench 4.0

A benchmark to measure and evolve with the frontier of agent work: real terminal environments, real software engineering tasks, and a rolling task set that is revised as agents catch up to it.

By Resolution Rate
1
Image
GPT-6 Astra (max) • Codex
58.2%
2
Image
Fable 5.1 (max) • Claude Code
57.9%
3
Image
Opus 5 (xhigh) • Claude Code
53.9%
Open Benchmarks Grants

Terminal-Bench-Science

A benchmark for evaluating AI agents on workflows from researchers’ own work. Scientists, not model developers or vendors, set the bar for scientific capability in AI.

By Resolution Rate
1
Image
Fable 5.1 (max) • Claude Code
40.0%
2
Image
Opus 5 (max) • Claude Code
30.0%
3
Image
GPT-5.6 Sol (max) • Codex
22.4%
Open Benchmarks Grants

Terminal-Bench 3.0

The next frontier benchmark for agent work. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.

By Resolution Rate
1
Image
Opus 5 (max) • mini-SWE-agent
42.7%
2
Image
GPT-5.6 Sol (max) • Codex
34.6%
3
Image
Fable 5 (max) • Claude Code
34.1%

Senior SWE-bench

A benchmark for evaluating coding agents on senior-level engineering work: building features from realistic instructions, investigating bugs that require runtime investigation, and shipping code that aligns to existing codebase conventions.

Tasteful Solve Rate
1
Image
Claude Fable 5.1
34.7%
2
Image
Claude Fable 5
34.7%
3
Image
Claude Opus 5
34.7%
Open Benchmarks Grants

OSWorld 2.0

A benchmark for evaluating computer-use agents on long-horizon, real-world workflows: 108 authentic tasks across 31 self-hosted web environments and professional desktop applications.

By Binary Accuracy (500 steps)
1
Image
Claude Opus 5 (max)
44.33%
2
Image
Claude Opus 5 (high)
36.89%
3
Image
Claude Opus 5 (xhigh)
33.33%
Open Benchmarks Grants

SlopCode Bench

Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.

Top Models by Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3-Codex
26.02%
3
Image
GPT-5.4
23.47%
Open Benchmarks Grants

Continual Learning Bench

Evaluates whether AI systems improve from prior experience across sequential, stateful tasks, measuring real in-context learning, not just raw capability.

Top Systems (Agg. Reward)
1
Image
ICL · Claude Sonnet 4.6
+0.196
2
Image
ICL · GPT-5.4
+0.189
3
Image
Claude Code · Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.