Open Benchmarks Grants
agent-le-logo

Agents' Last Exam

A benchmark for evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes. 55 sub-industries, 1,500+ tasks toward a 5,000-task target, sourced and validated by 300+ industry experts.

Built with
ImageImageSnorkel AI logo lockup mono white outline png
Overview

Agents’ Last Exam (ALE) is building the broadest-coverage agent evaluation benchmark to date, measuring performance on long-horizon, economically valuable tasks with verifiable outcomes. The benchmark covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy), spanning all 55 targeted sub-industries.

ALE-V1 ships 147 reference tasks across 55 industries as the current public subset of a 1,500+ task corpus. Many tasks require private data or licensed software and remain in a separate private pool. ALE uses rolling evaluation: every ~6 months a new public subset is published with fresh instances, while private tasks rotate in and retired public tasks rotate out, to limit benchmark leakage.

Leaderboard

Rank Harness Model Effort Pass Rate Score Est. Cost Runtime Input Tokens Output Tokens
1 Codex gpt-5-6-sol XHigh
30.6%
53.6%
$762 94h 39m 762.9M 3.8M
2 Codex gpt-5-6-sol High
30.6%
52.1%
$577 94h 33m 594.0M 2.9M
3 Codex gpt-5-6-sol Max
29.6%
52.6%
$1,085 103h 46m 1.1B 5.4M
4 Codex gpt-5-6-luna XHigh
29.6%
48.3%
$235 66h 7m 1.4B 4.2M
5 Codex gpt-5-6-sol Medium
29.5%
51.9%
$518 91h 17m 539.9M 2.2M
6 Codex gpt-5-6-luna Max
28.3%
49.2%
$391 66h 27m 2.5B 6.8M
7 Codex gpt-5-6-terra Max
28%
50.5%
$545 118h 39m 1.2B 5.8M
8 Codex gpt-5-6-terra XHigh
27.6%
48.5%
$381 87h 20m 822.5M 3.8M
9 Claude Code claude-opus-4-8 Max
27%
45.2%
$3,985 179h 19m 1.7B 22.1M
10 Codex gpt-5-5 XHigh
26.6%
47.9%
$602 97h 8m 560.5M 6.4M
11 Codex gpt-5-6-terra High
26%
46.3%
$264 98h 34m 533.9M 2.8M
12 Claude Code claude-fable-5 XHigh
25.7%
48.7%
$4,340 70h 0m 2.7B 10.4M
13 Codex gpt-5-5 High
24.2%
44%
$416 75h 15m 489.9M 3.5M
14 Codex gpt-5-6-luna High
23.7%
44.9%
$128 45h 53m 699.4M 2.6M
15 Codex gpt-5-6-sol Low
23.6%
44.8%
$248 75h 8m 252.0M 1.3M
16 ALE-Claw gpt-5-5 High
23%
45.8%
$310 35h 34m 334.3M 2.3M
17 Codex gpt-5-6-terra Medium
22.7%
42.6%
$128 86h 4m 244.6M 1.6M
18 Claude Code claude-opus-4-8 XHigh
22.4%
42.3%
$2,815 144h 16m 1.2B 16.5M
19 Claude Code claude-fable-5 Adaptive
22%
40.5%
$2,315 106h 14m 880.0M 9.4M
20 OpenClaw gpt-5-5 High
21.1%
41%
$447 82h 38m 469.3M 3.3M
21 Cursor CLI gpt-5-5 Medium
20.7%
39.6%
$174 58h 47m 148.1M 1.7M
22 OpenClaw gpt-5-4 High
20.5%
37.3%
$274 127h 46m 488.8M 7.3M
23 Claude Code claude-opus-4-8 High
20.4%
38.7%
$1,626 82h 49m 1.8B 6.4M
24 Cursor CLI claude-opus-4-7 High
20.4%
41.8%
$1,899 63h 49m 429.5M 4.5M
25 Claude Code glm-5-2 Max
20.4%
40.6%
$1,086 107h 42m 1.3B 7.7M
26 Codex gpt-5-6-terra Low
20.4%
40.2%
$89 76h 23m 171.7M 1.1M
27 Cursor CLI composer-2-5 Adaptive
20.4%
38.5%
$177 77h 48m 338.8M 2.9M
28 Claude Code seed-2.1-pro -
19.5%
41.4%
$936 155h 14m 2.1B 8.4M
29 Droid gpt-5-5 High
19.1%
38.6%
$242 63h 49m 234.8M 2.2M
30 ALE-Claw claude-opus-4-7 High
18.4%
40.5%
$1,132 77h 39m 1.3B 5.6M
31 Codex gpt-5-5 Medium
18.4%
39.8%
$358 75h 11m 294.4M 2.9M
32 Codex gpt-5-6-luna Medium
17.1%
37.5%
$57 42h 8m 287.2M 1.1M
33 Codex gpt-5-5 Low
17.1%
36.4%
$147 38h 40m 109.3M 1.3M
34 Gemini CLI gemini-3-1-pro High
15.8%
32%
$2,018 92h 21m 1.2B 3.5M
35 OpenClaw claude-opus-4-7 High
15.1%
34.6%
$1,689 124h 10m 807.3M 4.0M
36 OpenClaw gemini-3-1-pro High
14.1%
28.7%
$2,969 147h 6m 3.4B 3.7M
37 Claude Code claude-opus-4-7 High
13.2%
35.1%
$1,793 42h 36m 456.4M 3.7M
38 OpenClaw seed-2.1-pro High
13.1%
32.2%
$440 190h 18m 1.6B 20.4M
39 Droid claude-opus-4-7 High
12.8%
31%
$1,356 28h 16m 352.1M 2.7M
40 OpenClaw deepseek-v4-pro High
12.4%
27.6%
$273 130h 47m 511.7M 4.7M
41 OpenClaw qwen3-7-max High
11.8%
31.1%
$664 101h 0m 1.4B 17.6M
42 Codex gpt-5-6-luna Low
11.8%
30.5%
$27 24h 35m 119.9M 610.3K
43 ALE-Claw gpt-5-4 High
11.8%
28.2%
$335 55h 28m 1.1B 2.0M
44 OpenClaw glm-5-1 High
11.5%
28.1%
$440 168h 39m 755.8M 5.9M
45 OpenClaw kimi-k2-6 High
9.2%
21.7%
$123 183h 18m 295.9M 6.0M
46 OpenClaw qwen3-6-plus High
8.6%
24.3%
$254 134h 43m 738.9M 7.0M
47 OpenClaw mimo-v2-5 High
8.6%
23.6%
$43 107h 22m 394.9M 4.0M
48 Grok CLI grok-4-3 -
6.6%
20.1%
$285 43h 7m 223.1M 2.3M
49 OpenClaw minimax-m2-7 High
5.9%
14.2%
$27 144h 37m 277.2M 4.4M
50 OpenClaw grok-4-3 High
4.3%
15.5%
$80 79h 4m 168.7M 2.7M
Rank Harness Model Effort Pass Rate Score Est. Cost Runtime Input Tokens Output Tokens
1 Codex gpt-5-6-sol XHigh
28.6%
53.6%
$437 49h 46m 443.5M 2.5M
2 Codex gpt-5-6-sol Medium
28.6%
52.8%
$279 47h 32m 300.1M 1.4M
3 Codex gpt-5-6-luna XHigh
28.6%
49.1%
$144 24h 53m 910.7M 2.8M
4 Codex gpt-5-6-luna Max
27.6%
50.7%
$246 35h 22m 1.6B 4.3M
5 Codex gpt-5-6-sol High
27.1%
50.9%
$344 51h 12m 359.1M 1.9M
6 Codex gpt-5-5 XHigh
27.1%
49.2%
$385 47h 34m 357.1M 4.3M
7 Codex gpt-5-6-sol Max
26.7%
50.7%
$598 52h 15m 609.6M 3.5M
8 Codex gpt-5-6-terra XHigh
26.2%
49.4%
$235 49h 6m 528.0M 2.6M
9 Codex gpt-5-6-terra Max
25.7%
49.8%
$358 58h 41m 843.8M 3.9M
10 Claude Code claude-opus-4-8 Max
25.7%
44.1%
$1,953 110h 57m 642.2M 15.3M
11 ALE-Claw gpt-5-5 High
24.8%
48.1%
$109 16h 44m 77.0M 1.3M
12 Codex gpt-5-6-terra High
24.3%
46.6%
$146 46h 22m 297.4M 1.8M
13 Claude Code claude-fable-5 XHigh
23.8%
48.5%
$1,663 34h 35m 900.2M 6.4M
14 Codex gpt-5-6-luna High
23.8%
46%
$73 16h 49m 413.7M 1.7M
15 OpenClaw gpt-5-5 High
23.8%
44.4%
$219 31h 30m 211.2M 2.2M
16 Claude Code glm-5-2 Max
23.8%
43.4%
$981 64h 18m 1.2B 6.3M
17 Codex gpt-5-6-sol Low
23.3%
46.3%
$132 44h 41m 140.2M 820.3K
18 Codex gpt-5-5 High
23.3%
44.6%
$221 31h 18m 292.7M 2.1M
19 OpenClaw gpt-5-4 High
22.5%
41.8%
$169 50h 7m 285.8M 4.9M
20 Codex gpt-5-6-terra Medium
22.4%
43.9%
$65 53h 4m 125.1M 1.0M
21 Cursor CLI claude-opus-4-7 High
21.9%
44.7%
$1,045 24h 54m 207.5M 2.5M
22 Claude Code claude-fable-5 Adaptive
21.9%
44.6%
$880 69h 43m 301.4M 6.9M
23 Cursor CLI gpt-5-5 Medium
21.9%
41.8%
$97 14h 30m 68.5M 1.1M
24 Claude Code claude-opus-4-8 High
21.9%
42.7%
$1,016 58h 43m 1.1B 4.2M
25 Codex gpt-5-6-terra Low
20.5%
42.5%
$50 45h 45m 99.3M 779.0K
26 ALE-Claw claude-opus-4-7 High
20%
43.3%
$416 44h 15m 389.6M 2.9M
27 Claude Code claude-opus-4-8 XHigh
20%
42.1%
$1,713 84h 37m 641.0M 11.9M
28 Cursor CLI composer-2-5 Adaptive
19%
40.8%
$81 60h 10m 153.2M 1.8M
29 Droid gpt-5-5 High
19%
39.5%
$136 37h 27m 117.0M 1.5M
30 Claude Code seed-2.1-pro -
19%
41.4%
$602 88h 11m 1.1B 5.8M
31 Codex gpt-5-6-luna Medium
18.1%
39.8%
$25 13h 32m 127.3M 721.7K
32 Codex gpt-5-5 Low
18.1%
39.4%
$96 16h 26m 75.9M 901.9K
33 Codex gpt-5-5 Medium
17.1%
39.4%
$195 26h 28m 157.9M 1.9M
34 Gemini CLI gemini-3-1-pro High
17.1%
36.2%
$1,010 52h 27m 450.8M 2.4M
35 OpenClaw claude-opus-4-7 High
16.2%
37.8%
$960 53h 29m 516.2M 2.8M
36 OpenClaw gemini-3-1-pro High
15.7%
31.7%
$1,303 66h 23m 1.4B 2.3M
37 Claude Code claude-opus-4-7 High
14.3%
38%
$1,165 25h 47m 281.0M 2.7M
38 Droid claude-opus-4-7 High
14.3%
33.7%
$710 19h 17m 176.9M 1.7M
39 OpenClaw deepseek-v4-pro High
14.1%
30.8%
$161 56h 12m 312.8M 3.3M
40 Forgecode claude-sonnet-4-6 Medium
13.3%
28.5%
$100 51h 31m 136.5M 1.7M
41 OpenClaw qwen3-7-max High
13.3%
34%
$534 34h 23m 1.0B 15.8M
42 ALE-Claw gpt-5-4 High
13.3%
33.1%
$168 26h 31m 541.3M 1.1M
43 Hermes claude-sonnet-4-6 Medium
13.2%
32%
$437 25h 12m 195.6M 2.0M
44 OpenClaw glm-5-1 High
12.9%
30.8%
$213 73h 57m 361.0M 3.3M
45 Terminus 2 claude-sonnet-4-6 Off
11.9%
30.9%
$320 73h 1m 760.4M 3.1M
46 Codex gpt-5-6-luna Low
11.4%
32.8%
$13 9h 12m 56.3M 407.1K
47 OpenClaw claude-sonnet-4-6 High
11.4%
31%
$181 33h 51m 247.0M 1.9M
48 OpenClaw seed-2.1-pro High
11.4%
33.2%
$440 98h 0m 1.4B 18.4M
49 OpenClaw qwen3-6-plus High
10.5%
28.6%
$128 62h 28m 366.8M 4.4M
50 OpenClaw mimo-v2-5 High
10%
26.5%
$32 51h 24m 286.4M 3.0M
51 Openhands claude-sonnet-4-6 Not reported
9%
19.8%
$243 35h 47m 349.9M 4.4M
52 OpenClaw kimi-k2-6 High
8.1%
21.2%
$91 89h 31m 222.9M 4.4M
53 Grok CLI grok-4-3 -
7.6%
24.3%
$205 36h 12m 161.0M 1.7M
54 OpenClaw minimax-m2-7 High
5.7%
14.6%
$22 52h 56m 242.7M 3.2M
55 OpenClaw grok-4-3 High
4.3%
17.5%
$61 36h 8m 135.8M 2.1M
Rank Harness Model Effort Pass Rate Score Est. Cost Runtime Input Tokens Output Tokens
1 Codex gpt-5-6-sol Max
47.8%
78.8%
$322 40h 29m 304.2M 1.8M
2 Codex gpt-5-6-luna XHigh
47.8%
74.1%
$69 19h 22m 359.7M 1.6M
3 Codex gpt-5-6-sol High
47%
77.5%
$172 33h 30m 155.3M 1.0M
4 Codex gpt-5-6-sol XHigh
45.5%
78.3%
$210 34h 43m 190.0M 1.3M
5 Codex gpt-5-6-sol Medium
45.3%
75.6%
$137 34h 16m 124.2M 787.5K
6 Codex gpt-5-6-terra XHigh
44.8%
73.9%
$96 31h 27m 170.7M 1.3M
7 Codex gpt-5-5 XHigh
44%
72.4%
$170 29h 32m 126.2M 2.3M
8 Codex gpt-5-6-luna Max
43.3%
73.2%
$96 21h 49m 541.7M 2.2M
9 Codex gpt-5-6-terra High
43.3%
70.8%
$84 40h 16m 156.0M 1.0M
10 Claude Code claude-opus-4-8 Max
43.3%
64%
$1,106 49h 40m 310.5M 8.3M
11 Codex gpt-5-6-terra Max
41.8%
73.3%
$141 37h 13m 268.6M 2.0M
12 Codex gpt-5-6-sol Low
39.3%
66.8%
$80 26h 43m 71.0M 467.6K
13 Codex gpt-5-6-terra Medium
38.8%
66%
$47 36h 15m 84.6M 590.9K
14 Codex gpt-5-5 High
38.6%
66.2%
$121 27h 0m 162.7M 1.3M
15 Claude Code claude-fable-5 XHigh
37.3%
71.1%
$1,018 21h 14m 504.9M 3.6M
16 Codex gpt-5-6-luna High
35.8%
66.2%
$46 16h 50m 228.8M 1.1M
17 OpenClaw gpt-5-5 High
35.8%
65.7%
$218 41h 19m 223.0M 1.5M
18 Claude Code claude-opus-4-8 XHigh
35.8%
62.7%
$683 42h 8m 201.3M 6.0M
19 Claude Code claude-opus-4-8 High
34.3%
56.8%
$446 21h 15m 397.7M 2.2M
20 Claude Code claude-fable-5 Adaptive
34.3%
63.4%
$947 31h 49m 341.7M 3.1M
21 Cursor CLI composer-2-5 Adaptive
34.3%
61.1%
$50 22h 1m 94.1M 1.1M
22 OpenClaw gpt-5-4 High
33.6%
57.8%
$69 40h 56m 86.0M 2.4M
23 ALE-Claw gpt-5-5 High
32.8%
67.4%
$148 14h 57m 167.2M 1.0M
24 Claude Code glm-5-2 Max
32.8%
59.1%
$204 32h 45m 187.0M 2.7M
25 Cursor CLI gpt-5-5 Medium
32.1%
60.8%
$68 30h 10m 51.8M 741.1K
26 Codex gpt-5-6-terra Low
32.1%
60.7%
$34 32h 24m 62.2M 448.1K
27 Cursor CLI claude-opus-4-7 High
29.9%
61.2%
$558 33h 26m 110.3M 1.6M
28 Codex gpt-5-6-luna Medium
29.9%
59.7%
$19 18h 46m 90.6M 452.4K
29 Codex gpt-5-5 Medium
29.9%
59.3%
$110 13h 22m 81.7M 1.1M
30 Droid gpt-5-5 High
29.9%
58.2%
$106 23h 26m 98.4M 962.6K
31 ALE-Claw claude-opus-4-7 High
28.4%
60.5%
$312 20h 40m 359.2M 1.9M
32 Droid claude-opus-4-7 High
27.6%
60.2%
$738 16h 6m 176.2M 1.5M
33 Claude Code seed-2.1-pro -
26.9%
58.9%
$376 66h 30m 697.7M 3.1M
34 OpenClaw claude-opus-4-7 High
26.9%
56.5%
$508 47h 42m 201.7M 1.5M
35 Gemini CLI gemini-3-1-pro High
26.9%
53.5%
$342 27h 40m 239.4M 1.1M
36 OpenClaw gemini-3-1-pro High
26.1%
48.3%
$575 48h 37m 616.7M 938.9K
37 Codex gpt-5-5 Low
25.4%
53.5%
$48 10h 16m 28.2M 528.4K
38 Codex gpt-5-6-luna Low
22.4%
47.1%
$11 14h 7m 46.7M 263.7K
39 Claude Code claude-opus-4-7 High
20.9%
54.3%
$496 16h 17m 120.4M 1.1M
40 ALE-Claw gpt-5-4 High
20.9%
44.3%
$66 17h 22m 178.6M 684.8K
41 OpenClaw glm-5-1 High
20.1%
45.6%
$108 62h 17m 183.5M 1.6M
42 OpenClaw deepseek-v4-pro High
19.9%
43.8%
$109 58h 50m 208.2M 1.9M
43 OpenClaw qwen3-7-max High
17.9%
46.9%
$247 38h 11m 502.0M 7.3M
44 OpenClaw seed-2.1-pro High
17.9%
46.9%
$201 76h 7m 640.9M 8.9M
45 OpenClaw kimi-k2-6 High
15.7%
35.1%
$47 61h 6m 118.5M 2.1M
46 OpenClaw qwen3-6-plus High
12.7%
35.8%
$90 60h 13m 259.7M 2.7M
47 OpenClaw mimo-v2-5 High
11.9%
35.1%
$13 43h 37m 105.4M 1.5M
48 OpenClaw minimax-m2-7 High
10.4%
24.5%
$9 63h 22m 98.2M 1.4M
49 Grok CLI grok-4-3 -
9%
30.4%
$127 21h 59m 99.8M 1.0M
50 OpenClaw grok-4-3 High
6.7%
24.4%
$35 31h 15m 73.8M 1.2M
Rank Harness Model Effort Pass Rate Score Est. Cost Runtime Input Tokens Output Tokens
1 Codex gpt-5-6-sol XHigh
30%
46.6%
$247 28h 47m 235.4M 1.3M
2 Codex gpt-5-6-terra Max
27.3%
45.6%
$150 38h 15m 316.9M 1.8M
3 Codex gpt-5-6-sol Medium
27.3%
43.9%
$157 29h 58m 142.1M 732.3K
4 Codex gpt-5-6-sol High
27.3%
42.4%
$166 29h 53m 153.0M 942.9K
5 Codex gpt-5-6-luna Max
27.3%
40.7%
$107 19h 52m 652.4M 2.2M
6 Codex gpt-5-6-luna XHigh
27.3%
40.4%
$61 22h 35m 351.8M 1.4M
7 Codex gpt-5-6-sol Max
25.5%
44.7%
$300 28h 9m 286.3M 1.8M
8 Claude Code claude-fable-5 XHigh
25.5%
42%
$1,826 25h 34m 1.1B 3.8M
9 Codex gpt-5-6-luna High
25.5%
41.9%
$38 10h 25m 202.0M 828.7K
10 Codex gpt-5-6-terra XHigh
24.5%
41.3%
$108 29h 0m 215.8M 1.2M
11 Claude Code seed-2.1-pro -
24.1%
40.1%
$289 42h 44m 639.5M 3.2M
12 Claude Code claude-opus-4-8 Max
23.6%
43.1%
$1,324 59h 37m 524.3M 8.2M
13 ALE-Claw gpt-5-5 High
23.6%
41.1%
$70 8h 28m 54.9M 734.8K
14 Codex gpt-5-5 High
22.4%
36.7%
$130 20h 57m 130.8M 1.2M
15 Codex gpt-5-5 XHigh
21.8%
38.7%
$178 29h 36m 142.8M 2.2M
16 Codex gpt-5-6-terra High
20.9%
38.3%
$85 36h 27m 164.1M 940.7K
17 Claude Code claude-fable-5 Adaptive
20.9%
34.1%
$608 32h 11m 327.7M 3.7M
18 Cursor CLI claude-opus-4-7 High
20%
37.6%
$842 14h 48m 215.6M 1.9M
19 Cursor CLI gpt-5-5 Medium
20%
32.7%
$50 17h 24m 36.4M 552.1K
20 Claude Code claude-opus-4-8 XHigh
20%
38.3%
$856 43h 56m 300.4M 5.6M
21 OpenClaw gpt-5-4 High
19.4%
34.3%
$123 57h 48m 239.6M 2.9M
22 Codex gpt-5-6-sol Low
19.1%
38.3%
$77 23h 19m 73.1M 437.4K
23 Claude Code glm-5-2 Max
18.2%
37%
$390 35h 45m 338.8M 2.7M
24 ALE-Claw claude-opus-4-7 High
18.2%
36.6%
$246 15h 31m 288.8M 1.7M
25 Codex gpt-5-6-terra Low
18.2%
34.5%
$31 22h 57m 54.5M 398.1K
26 Codex gpt-5-6-terra Medium
18.2%
34%
$44 28h 6m 78.9M 555.2K
27 OpenClaw gpt-5-5 High
18.2%
32.1%
$101 22h 58m 98.6M 1.1M
28 Cursor CLI composer-2-5 Adaptive
18.2%
30.8%
$68 25h 41m 130.0M 1.1M
29 Claude Code claude-opus-4-8 High
18.2%
37.8%
$569 26h 8m 698.1M 2.3M
30 Droid gpt-5-5 High
16.4%
33.4%
$69 15h 37m 58.6M 792.2K
31 Codex gpt-5-5 Medium
16.4%
33.3%
$128 26h 53m 102.4M 1.1M
32 Codex gpt-5-5 Low
16.4%
32.4%
$45 14h 49m 29.5M 425.3K
33 OpenClaw seed-2.1-pro High
13.5%
23.6%
$171 46h 53m 572.9M 7.7M
34 Claude Code claude-opus-4-7 High
12.7%
29.1%
$747 12h 5m 202.0M 1.5M
35 Codex gpt-5-6-luna Medium
12.7%
28.8%
$18 8h 56m 85.6M 355.6K
36 Gemini CLI gemini-3-1-pro High
12.7%
26.4%
$962 25h 34m 481.7M 1.7M
37 OpenClaw claude-opus-4-7 High
10.9%
27.5%
$370 29h 47m 163.0M 1.1M
38 OpenClaw qwen3-7-max High
10.9%
27.2%
$280 35h 46m 585.1M 7.1M
39 OpenClaw deepseek-v4-pro High
10.9%
23.8%
$68 38h 12m 114.5M 1.5M
40 OpenClaw gemini-3-1-pro High
10.9%
23.6%
$1,392 62h 47m 1.6B 1.7M
41 ALE-Claw gpt-5-4 High
9.1%
22.9%
$179 22h 36m 595.7M 930.0K
42 OpenClaw glm-5-1 High
9.1%
21.7%
$165 61h 57m 282.4M 2.2M
43 OpenClaw mimo-v2-5 High
9.1%
20.8%
$17 34h 51m 159.9M 1.4M
44 OpenClaw qwen3-6-plus High
8.2%
22.9%
$111 46h 37m 322.1M 3.1M
45 Codex gpt-5-6-luna Low
7.3%
24.9%
$9 5h 32m 36.3M 207.5K
46 Grok CLI grok-4-3 -
7.3%
17%
$95 10h 30m 74.0M 896.2K
47 OpenClaw kimi-k2-6 High
6.4%
18.2%
$41 73h 35m 93.6M 2.4M
48 OpenClaw grok-4-3 High
3.6%
12.9%
$25 24h 45m 51.6M 1.0M
49 Droid claude-opus-4-7 High
3.6%
10.9%
$167 3h 8m 39.9M 472.1K
50 OpenClaw minimax-m2-7 High
3.6%
8.4%
$10 46h 17m 95.2M 1.7M
Rank Harness Model Effort Pass Rate Score Est. Cost Runtime Input Tokens Output Tokens
1 Claude Code claude-fable-5 XHigh
7.9%
16.5%
$2,091 30h 16m 1.6B 3.8M
2 Codex gpt-5-6-sol XHigh
5.3%
19.4%
$325 38h 53m 352.3M 1.4M
3 Codex gpt-5-6-sol High
5.3%
18.7%
$253 39h 4m 297.5M 1.1M
4 Codex gpt-5-6-sol Medium
3.9%
19.4%
$250 32h 36m 294.9M 783.9K
5 Codex gpt-5-6-sol Max
2.6%
16.4%
$488 42h 49m 523.8M 2.1M
6 Codex gpt-5-6-luna Max
2.6%
15%
$197 31h 35m 1.4B 2.6M
7 Codex gpt-5-6-terra Max
2.6%
14.9%
$264 50h 55m 678.1M 2.2M
8 Codex gpt-5-5 Medium
2.6%
14.2%
$167 41h 36m 153.9M 905.0K
9 Claude Code glm-5-2 Max
2.6%
12.9%
$638 47h 56m 782.7M 2.7M
10 ALE-Claw gpt-5-5 High
2.6%
12.8%
$107 13h 44m 126.3M 704.4K
11 Cursor CLI claude-opus-4-7 High
2.6%
11.4%
$588 18h 31m 121.0M 1.2M
12 Droid gpt-5-5 High
2.6%
11.3%
$76 26h 27m 83.6M 551.8K
13 Cursor CLI gpt-5-5 Medium
2.6%
10.7%
$61 17h 12m 63.9M 456.4K
14 Claude Code claude-opus-4-8 XHigh
2.6%
10.4%
$1,395 70h 29m 739.8M 5.9M
15 Claude Code claude-opus-4-8 Max
2.6%
14.4%
$1,805 84h 44m 961.0M 7.3M
16 Codex gpt-5-5 XHigh
1.3%
16.1%
$283 42h 39m 316.9M 2.2M
17 Codex gpt-5-5 High
0%
13.4%
$188 33h 14m 216.4M 1.2M
18 Codex gpt-5-6-sol Low
0%
12.6%
$96 30h 32m 111.3M 408.6K
19 Codex gpt-5-6-terra Medium
0%
12.6%
$43 25h 34m 90.6M 486.7K
20 Codex gpt-5-6-terra High
0%
12%
$106 28h 23m 236.9M 967.5K
21 Codex gpt-5-6-terra Low
0%
11.2%
$31 23h 59m 65.9M 335.9K
22 Codex gpt-5-6-luna High
0%
11%
$47 20h 58m 284.9M 772.0K
23 Codex gpt-5-6-luna XHigh
0%
11%
$110 30h 12m 742.9M 1.3M
24 Codex gpt-5-6-terra XHigh
0%
10.9%
$186 33h 22m 449.2M 1.4M
25 OpenClaw gpt-5-5 High
0%
10.9%
$155 24h 52m 178.2M 997.0K
26 Codex gpt-5-6-luna Medium
0%
10.4%
$24 15h 42m 132.8M 319.4K
27 Claude Code claude-opus-4-7 High
0%
10%
$685 17h 6m 170.1M 1.4M
28 Claude Code seed-2.1-pro -
0%
10.8%
$319 52h 3m 814.8M 2.8M
29 Codex gpt-5-5 Low
0%
9.9%
$64 16h 3m 59.4M 378.0K
30 OpenClaw seed-2.1-pro High
0%
10.4%
$85 75h 30m 455.8M 4.8M
31 Codex gpt-5-6-luna Low
0%
8.9%
$9 6h 13m 42.6M 162.5K
32 Cursor CLI composer-2-5 Adaptive
0%
8.8%
$78 34h 44m 151.0M 948.3K
33 ALE-Claw gpt-5-4 High
0%
8.1%
$100 17h 39m 321.4M 540.7K
34 Droid claude-opus-4-7 High
0%
8%
$496 10h 18m 145.7M 920.1K
35 ALE-Claw claude-opus-4-7 High
0%
7.9%
$594 43h 8m 707.3M 2.2M
36 OpenClaw gpt-5-4 High
0%
7.3%
$107 36h 16m 206.9M 2.5M
37 OpenClaw qwen3-7-max High
0%
6.4%
$173 33h 3m 377.8M 4.1M
38 OpenClaw glm-5-1 High
0%
6.2%
$204 57h 10m 354.0M 2.5M
39 Claude Code claude-fable-5 Adaptive
0%
5.2%
$838 52h 15m 231.2M 3.3M
40 Claude Code claude-opus-4-8 High
0%
7%
$735 43h 8m 841.1M 2.2M
41 OpenClaw qwen3-6-plus High
0%
5%
$78 35h 46m 228.9M 1.9M
42 OpenClaw claude-opus-4-7 High
0%
4.3%
$842 50h 0m 452.9M 1.5M
43 OpenClaw mimo-v2-5 High
0%
3.4%
$15 33h 3m 135.8M 1.3M
44 OpenClaw gemini-3-1-pro High
0%
3.1%
$1,200 45h 9m 1.4B 1.1M
45 OpenClaw deepseek-v4-pro High
0%
2.5%
$108 39h 7m 208.8M 1.7M
46 Grok CLI grok-4-3 -
0%
2.3%
$71 11h 35m 56.2M 477.3K
47 OpenClaw grok-4-3 High
0%
2.3%
$22 25h 4m 47.2M 621.3K
48 OpenClaw kimi-k2-6 High
0%
1.6%
$39 60h 46m 88.1M 1.8M
49 OpenClaw minimax-m2-7 High
0%
1.3%
$9 44h 14m 87.7M 1.5M
50 Gemini CLI gemini-3-1-pro High
0%
0.9%
$733 40h 2m 497.3M 821.0K

*claude-fable-5: the variant Anthropic served during evaluation may differ from the published model's full capability tier, and re-runs cannot guarantee the higher-tier variant is selected. These numbers may understate the model's true ceiling. Learn more.

Pass rate vs input tokens

Efficiency frontier for this evaluation split. Up and to the left is better (higher pass rate at lower input-token spend). Several harness/model pairs reach the top of the split at a fraction of the cost of others; some pairs spend an order of magnitude more tokens without a corresponding pass-rate gain.

Hover any dot for the harness, model, and exact metrics. 

Industry coverage

Six representative task families across the 55 sub-industries ALE covers.
Image

Motion & VFX

Animation and visual effects production tasks in Adobe After Effects.
Image

3D modeling

3D model creation and editing tasks in Siemens NX.
Image

Game development

Scene setup, asset placement, and rendering tasks in Unreal Engine.
Image

Mold flow analysis

Simulation and mold flow analysis tasks in Moldex3D manufacturing software.
Image

Architectural modeling

3D modeling and energy analysis workflows in Rhino 3D for urban design.
Image

Brain imaging

Neuroimaging analysis and brain structure segmentation tasks in FSLeyes.

Sample tasks

A selection from the 147 public ALE-V1 tasks across 14 task categories. Each task ships with a sandboxed environment, a hidden reference, and a deterministic grader. Slugs link to the task source.

business finance

sec_10k_financial_parsing

Parse a SEC 10-K filing into a structured financial schema. Multi-step extraction, table normalization, and cross-reference validation against the original document.

business finance

financial_stmt_reconstruction_aapl_fy2024
Reconstruct Apple’s FY2024 financial statement from primary disclosure documents. Validates whether the agent surfaces the exact reported figures and footnote-relevant adjustments.
engineering
mold-flow / 220089
Set up a Moldex3D mold-flow simulation, run it to convergence, and report fill time / pressure metrics matching the held-out reference run.

health medicine

Clinical_Variant_Annotation

Annotate a clinical variant set using standard pipelines (VEP, ClinVar, etc.) and produce a report graded against a curated reference.

life sciences

WGS_Variant_Calling
Run a whole-genome sequencing variant-calling pipeline and produce VCF output. Scored on precision and recall against a held-out truth VCF.

computing math

k8s_payment_api_root_cause_analysis
Diagnose a failing payment API in a Kubernetes cluster. Multi-hop investigation across logs, metrics, manifests, and traces, scored on the correct root-cause identification.

visual media

video_storyboard_001
Build a shot-by-shot video storyboard from a brief, formatted to industry conventions. Graded on coverage, continuity, and adherence to the reference shot list.
legal
legal_dr_fees_01
Compute legal fees from a billing register according to jurisdictional rules. Tests structured extraction plus rule-following against an authoritative reference total.

Methodology

Metrics

Pass Rate — fraction of tasks the agent fully completed (strict success). Score — average graded outcome across all tasks, including partial credit. Both computed by deterministic graders against hidden references.

Verifiable Outcomes
Hidden references plus deterministic graders, not LLM-as-a-judge. Tasks sourced from real professional workflows (After Effects, Siemens NX, Unreal Engine, Moldex3D, Rhino 3D, FSLeyes, and 49 more applications) and validated by domain experts before inclusion.
Rolling Evaluation

Every ~6 months, a new public subset releases with fresh instances. Private tasks rotate into the public pool, retired public tasks rotate out, and held-out private tasks score the official leaderboard, to limit benchmark leakage.

Reference Harnesses

Two open harnesses ship with the framework: the official Claude Code CLI and the in-tree OpenClaw harness. Submissions also include Codex, Cursor CLI, Droid, Gemini CLI, Grok CLI, and the ALE Claw reference harness.

Acknowledgments

Agents’ Last Exam is co-led by UC Berkeley RDI and the RDI Foundation, with funding support and contributions from Snorkel AI via the Open Benchmarks Grants program. The benchmark draws task contributions from 300+ industry experts across 44 academic institutions (MIT, Harvard, Stanford, UC Berkeley, Oxford, CMU, Caltech, ETH Zurich, Yale, Columbia, and more) and industry organizations including Goldman Sachs, JPMorgan, Morgan Stanley, PIMCO, Meta, Amazon, Adobe, Oracle, Hippocratic AI, and HubSpot.

Advisory Committee includes George Em Karniadakis (Brown), Tapio Schneider (Caltech), Teresa Head-Gordon (UC Berkeley), Laure Zanna (NYU), Jack Gallant (UC Berkeley), Tarek Zohdi (UC Berkeley), Ida Sim (UCSF), Arvind Rao (U Michigan), Kaan Ozbay (NYU), Carl Boettiger (UC Berkeley), Kyle Steinfeld (UC Berkeley), Yamini Rangan (HubSpot), and Bradley Rothenberg (nTop).

FAQs

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

Senior SWE-bench

A benchmark for evaluating coding agents on senior-level engineering work: building features from realistic instructions, investigating bugs that require runtime investigation, and shipping code that aligns to existing codebase conventions.

Tasteful Solve Rate
1
Image
Claude Fable 5
29.1%
2
Image
Claude Opus 4.8
25.0%
3
Image
GPT-5.6 Sol
24.4%
Open Benchmarks Grants

OSWorld 2.0

Long-horizon professional workflows with verifiable outcomes across 55 sub-industries. 147 public tasks of a 1,500+ task corpus, sourced and validated by 300+ industry experts.

By binary accuracy (300 steps)
1
Image
gpt-5-5 · xhigh
13%
2
Image
Claude Opus 4.7 · max
13%
3
Image
Claude Sonnet 4.6 · medium
8.3%

Agentic Coding

A benchmark for evaluating AI models on complex, real-world coding tasks that require multi-step reasoning, tool use, and autonomous problem-solving.

Top Models
1
Image
Claude Opus 4.6
65.2%
2
Image
Claude Opus 4.5
58.0%
3
Image
Claude Sonnet 4.5
57.6%
Open Benchmarks Grants

SlopCode Bench

Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.

Top Models by Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3-Codex
26.02%
3
Image
GPT-5.4
23.47%
Open Benchmarks Grants

Continual Learning Bench

Evaluates whether AI systems improve from prior experience across sequential, stateful tasks, measuring real in-context learning, not just raw capability.

Top Systems (Agg. Reward)
1
Image
ICL · Claude Sonnet 4.6
+0.196
2
Image
ICL · GPT-5.4
+0.189
3
Image
Claude Code · Sonnet 4.6
+0.185
Open Benchmarks Grants

Terminal-Bench 2.1

Terminal agent evaluation led by Stanford University and Laude Institute. v2.1 fixes 28 tasks from 2.0 and introduces continuous validation.

Top Submissions
1
Image
Codex CLI · GPT-5.5
83.4%
2
Image
Claude Code · Claude 5 Fable
83.1%
3
Image
Terminus 2 · Claude 5 Fable
80.4%
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.