Open Benchmarks Grants
agent-le-logo

Agents' Last Exam

A benchmark for evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes. 55 sub-industries, 1,500+ tasks toward a 5,000-task target, sourced and validated by 300+ industry experts.

Built with
ImageImageSnorkel AI logo lockup mono white outline png
Overview

Agents’ Last Exam (ALE) is building the broadest-coverage agent evaluation benchmark to date, measuring performance on long-horizon, economically valuable tasks with verifiable outcomes. The benchmark covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy), spanning all 55 targeted sub-industries.

ALE-V1 ships 147 reference tasks across 55 industries as the current public subset of a 1,500+ task corpus. Many tasks require private data or licensed software and remain in a separate private pool. ALE uses rolling evaluation: every ~6 months a new public subset is published with fresh instances, while private tasks rotate in and retired public tasks rotate out, to limit benchmark leakage.

Leaderboard

Rank Harness Model Effort Pass Rate Score Est. Cost Runtime Input Tokens Output Tokens
1 Codex GPT-5.6 Sol
30.6%
53.6%
$762 94h 39m 762.9M 3.8M
2 Codex GPT-5.6 Luna
30.3%
49.4%
$235 66h 7m 1.4B 4.2M
3 Kimi Code Kimi K3 Max
28.3%
51.6%
$606 186h 5m 1.4B 7.5M
4 Codex GPT-5.6 Terra
28%
50.7%
$545 118h 39m 1.2B 5.8M
5 Claude Code Qwen3 8 XHigh
27%
52.5%
$486 202h 41m 1.2B 12.4M
6 Grok Build Grok 4.5 High
27%
51.2%
$337 89h 8m 668.3M 5.4M
7 Claude Code Kimi K3 Max
27%
50.7%
$464 215h 32m 950.6M 7.4M
8 Claude Code Claude Opus 4.8
27%
45.1%
$3,985 179h 19m 1.7B 22.1M
9 Codex GPT-5.5
26.6%
47.9%
$602 97h 8m 560.5M 6.4M
10 Claude Code Claude Fable 5
25.7%
48.7%
$4,340 70h 0m 2.7B 10.4M
11 ALE-Claw GPT-5.5 High
23%
45.8%
$310 35h 34m 334.3M 2.3M
12 OpenClaw GPT-5.5 High
21.7%
42.1%
$447 82h 38m 469.3M 3.3M
13 Cursor CLI GPT-5.5 Medium
21.4%
40.7%
$174 58h 47m 148.1M 1.7M
14 Cursor CLI Claude Opus 4.7 High
21.1%
42.5%
$1,899 63h 49m 429.5M 4.5M
15 OpenClaw GPT-5.4 High
20.5%
37.3%
$274 127h 46m 488.8M 7.3M
16 Claude Code GLM-5.2 Max
20.4%
41.1%
$1,086 107h 42m 1.3B 7.7M
17 Cursor CLI Composer 2.5 Adaptive
20.4%
39.5%
$177 77h 48m 338.8M 2.9M
18 Droid GPT-5.5 High
19.7%
39.2%
$242 63h 49m 234.8M 2.2M
19 Claude Code Seed 2.1 Pro -
19.5%
41.4%
$936 155h 14m 2.1B 8.4M
20 ALE-Claw Claude Opus 4.7 High
18.4%
40.5%
$1,132 77h 39m 1.3B 5.6M
21 Gemini CLI Gemini 3.1 Pro High
16.4%
32.7%
$2,018 92h 21m 1.2B 3.5M
22 OpenClaw Claude Opus 4.7 High
15.8%
35.2%
$1,689 124h 10m 807.3M 4.0M
23 OpenClaw Gemini 3.1 Pro High
14.1%
29.3%
$2,969 147h 6m 3.4B 3.7M
24 Claude Code Claude Opus 4.7 High
13.8%
35.8%
$1,793 42h 36m 456.4M 3.7M
25 Droid Claude Opus 4.7 High
13.5%
31.7%
$1,356 28h 16m 352.1M 2.7M
26 OpenClaw Seed 2.1 Pro High
13.1%
32.2%
$440 190h 18m 1.6B 20.4M
27 OpenClaw DeepSeek V4 Pro High
12.4%
27.6%
$273 130h 47m 511.7M 4.7M
28 OpenClaw Qwen3 7-Max High
11.8%
31.6%
$664 101h 0m 1.4B 17.6M
29 ALE-Claw GPT-5.4 High
11.8%
28.6%
$335 55h 28m 1.1B 2.0M
30 OpenClaw GLM-5.1 High
11.5%
28.1%
$440 168h 39m 755.8M 5.9M
31 OpenClaw Kimi K2.6 High
9.2%
21.7%
$123 183h 18m 295.9M 6.0M
32 OpenClaw Qwen3 6-Plus High
8.6%
24.6%
$254 134h 43m 738.9M 7.0M
33 OpenClaw MiMo V2.5 High
8.6%
23.6%
$43 107h 22m 394.9M 4.0M
34 Grok CLI Grok 4.3 -
7.2%
21.3%
$285 43h 7m 223.1M 2.3M
35 OpenClaw MiniMax M2.7 High
5.9%
14.2%
$27 144h 37m 277.2M 4.4M
36 OpenClaw Grok 4.3 High
4.3%
15.8%
$80 79h 4m 168.7M 2.7M
Rank Harness Model Effort Pass Rate Score Est. Cost Runtime Input Tokens Output Tokens
1 Codex GPT-5.6 Luna
29.5%
50.7%
$144 24h 53m 910.7M 2.8M
2 Codex GPT-5.6 Sol
28.6%
53.6%
$437 49h 46m 443.5M 2.5M
3 Claude Code Kimi K3 Max
27.6%
52.8%
$251 109h 24m 465.2M 4.7M
4 Kimi Code Kimi K3 Max
27.6%
52%
$305 102h 15m 604.4M 5.0M
5 Codex GPT-5.5
27.1%
49.2%
$385 47h 34m 357.1M 4.3M
6 Codex GPT-5.6 Terra
26.2%
49.4%
$235 49h 6m 528.0M 2.6M
7 Grok Build Grok 4.5 High
25.7%
50.6%
$139 51h 49m 247.6M 3.3M
8 Claude Code Claude Opus 4.8
25.7%
44.1%
$1,953 110h 57m 642.2M 15.3M
9 Claude Code Qwen3 8 XHigh
24.8%
51.1%
$249 113h 20m 559.9M 6.8M
10 ALE-Claw GPT-5.5 High
24.8%
48.1%
$109 16h 44m 77.0M 1.3M
11 OpenClaw GPT-5.5 High
24.8%
46.1%
$219 31h 30m 211.2M 2.2M
12 Claude Code Claude Fable 5
23.8%
48.5%
$1,663 34h 35m 900.2M 6.4M
13 Claude Code GLM-5.2 Max
23.8%
44.2%
$981 64h 18m 1.2B 6.3M
14 Cursor CLI Claude Opus 4.7 High
22.9%
45.7%
$1,045 24h 54m 207.5M 2.5M
15 Cursor CLI GPT-5.5 Medium
22.9%
43.4%
$97 14h 30m 68.5M 1.1M
16 OpenClaw GPT-5.4 High
22.5%
41.8%
$169 50h 7m 285.8M 4.9M
17 ALE-Claw Claude Opus 4.7 High
20%
43.3%
$416 44h 15m 389.6M 2.9M
18 Droid GPT-5.5 High
20%
40.5%
$136 37h 27m 117.0M 1.5M
19 Cursor CLI Composer 2.5 Adaptive
19%
42.2%
$81 60h 10m 153.2M 1.8M
20 Claude Code Seed 2.1 Pro -
19%
41.4%
$602 88h 11m 1.1B 5.8M
21 Gemini CLI Gemini 3.1 Pro High
18.1%
37.2%
$1,010 52h 27m 450.8M 2.4M
22 OpenClaw Claude Opus 4.7 High
17.1%
38.7%
$960 53h 29m 516.2M 2.8M
23 OpenClaw Gemini 3.1 Pro High
15.7%
32.4%
$1,303 66h 23m 1.4B 2.3M
24 Claude Code Claude Opus 4.7 High
15.2%
38.9%
$1,165 25h 47m 281.0M 2.7M
25 Droid Claude Opus 4.7 High
15.2%
34.7%
$710 19h 17m 176.9M 1.7M
26 OpenClaw DeepSeek V4 Pro High
14.1%
30.8%
$161 56h 12m 312.8M 3.3M
27 OpenClaw Qwen3 7-Max High
13.3%
34.6%
$534 34h 23m 1.0B 15.8M
28 ALE-Claw GPT-5.4 High
13.3%
33.8%
$168 26h 31m 541.3M 1.1M
29 Forgecode Claude Sonnet 4.6 Medium
13.3%
29%
$100 51h 31m 136.5M 1.7M
30 Hermes Claude Sonnet 4.6 Medium
13.2%
32%
$437 25h 12m 195.6M 2.0M
31 OpenClaw GLM-5.1 High
12.9%
30.8%
$213 73h 57m 361.0M 3.3M
32 Terminus 2 Claude Sonnet 4.6 Off
11.9%
30.9%
$320 73h 1m 760.4M 3.1M
33 OpenClaw Seed 2.1 Pro High
11.4%
33.2%
$440 98h 0m 1.4B 18.4M
34 OpenClaw Claude Sonnet 4.6 High
11.4%
31.6%
$181 33h 51m 247.0M 1.9M
35 OpenClaw Qwen3 6-Plus High
10.5%
29%
$128 62h 28m 366.8M 4.4M
36 OpenClaw MiMo V2.5 High
10%
26.5%
$32 51h 24m 286.4M 3.0M
37 Openhands Claude Sonnet 4.6 Not reported
9%
19.8%
$243 35h 47m 349.9M 4.4M
38 Grok CLI Grok 4.3 -
8.6%
25.9%
$205 36h 12m 161.0M 1.7M
39 OpenClaw Kimi K2.6 High
8.1%
21.2%
$91 89h 31m 222.9M 4.4M
40 OpenClaw MiniMax M2.7 High
5.7%
14.6%
$22 52h 56m 242.7M 3.2M
41 OpenClaw Grok 4.3 High
4.3%
17.8%
$61 36h 8m 135.8M 2.1M
Rank Harness Model Effort Pass Rate Score Est. Cost Runtime Input Tokens Output Tokens
1 Codex GPT-5.6 Luna
49.3%
75.6%
$69 19h 22m 359.7M 1.6M
2 Codex GPT-5.6 Sol
47.8%
78.8%
$322 40h 29m 304.2M 1.8M
3 Codex GPT-5.6 Terra
44.8%
73.9%
$96 31h 27m 170.7M 1.3M
4 Codex GPT-5.5
44%
72.4%
$170 29h 32m 126.2M 2.3M
5 Claude Code Claude Opus 4.8
43.3%
64%
$1,106 49h 40m 310.5M 8.3M
6 Claude Code Kimi K3 Max
40.3%
71.6%
$145 66h 36m 270.2M 2.7M
7 Claude Code Qwen3 8 XHigh
38.8%
72.1%
$137 62h 29m 322.0M 4.7M
8 Grok Build Grok 4.5 High
38.8%
71.9%
$121 20h 54m 239.6M 2.0M
9 Kimi Code Kimi K3 Max
38.8%
69.7%
$129 50h 48m 256.7M 2.2M
10 Claude Code Claude Fable 5
37.3%
71.1%
$1,018 21h 14m 504.9M 3.6M
11 OpenClaw GPT-5.5 High
37.3%
67.2%
$218 41h 19m 223.0M 1.5M
12 Cursor CLI Composer 2.5 Adaptive
34.3%
62.3%
$50 22h 1m 94.1M 1.1M
13 Cursor CLI GPT-5.5 Medium
33.6%
62.3%
$68 30h 10m 51.8M 741.1K
14 OpenClaw GPT-5.4 High
33.6%
57.8%
$69 40h 56m 86.0M 2.4M
15 ALE-Claw GPT-5.5 High
32.8%
67.4%
$148 14h 57m 167.2M 1.0M
16 Claude Code GLM-5.2 Max
32.8%
60.3%
$204 32h 45m 187.0M 2.7M
17 Cursor CLI Claude Opus 4.7 High
31.3%
62.7%
$558 33h 26m 110.3M 1.6M
18 Droid GPT-5.5 High
29.9%
58.2%
$106 23h 26m 98.4M 962.6K
19 Droid Claude Opus 4.7 High
29.1%
61.7%
$738 16h 6m 176.2M 1.5M
20 ALE-Claw Claude Opus 4.7 High
28.4%
60.5%
$312 20h 40m 359.2M 1.9M
21 OpenClaw Claude Opus 4.7 High
28.4%
58%
$508 47h 42m 201.7M 1.5M
22 Gemini CLI Gemini 3.1 Pro High
28.4%
54.9%
$342 27h 40m 239.4M 1.1M
23 Claude Code Seed 2.1 Pro -
26.9%
58.9%
$376 66h 30m 697.7M 3.1M
24 OpenClaw Gemini 3.1 Pro High
26.1%
49.5%
$575 48h 37m 616.7M 938.9K
25 Claude Code Claude Opus 4.7 High
22.4%
55.8%
$496 16h 17m 120.4M 1.1M
26 ALE-Claw GPT-5.4 High
20.9%
45.4%
$66 17h 22m 178.6M 684.8K
27 OpenClaw GLM-5.1 High
20.1%
45.6%
$108 62h 17m 183.5M 1.6M
28 OpenClaw DeepSeek V4 Pro High
19.9%
43.8%
$109 58h 50m 208.2M 1.9M
29 OpenClaw Qwen3 7-Max High
17.9%
46.9%
$247 38h 11m 502.0M 7.3M
30 OpenClaw Seed 2.1 Pro High
17.9%
46.9%
$201 76h 7m 640.9M 8.9M
31 OpenClaw Kimi K2.6 High
15.7%
35.1%
$47 61h 6m 118.5M 2.1M
32 OpenClaw Qwen3 6-Plus High
12.7%
36.4%
$90 60h 13m 259.7M 2.7M
33 OpenClaw MiMo V2.5 High
11.9%
35.1%
$13 43h 37m 105.4M 1.5M
34 Grok CLI Grok 4.3 -
10.4%
31.8%
$127 21h 59m 99.8M 1.0M
35 OpenClaw MiniMax M2.7 High
10.4%
24.5%
$9 63h 22m 98.2M 1.4M
36 OpenClaw Grok 4.3 High
6.7%
24.4%
$35 31h 15m 73.8M 1.2M
Rank Harness Model Effort Pass Rate Score Est. Cost Runtime Input Tokens Output Tokens
1 Codex GPT-5.6 Sol
30%
46.6%
$247 28h 47m 235.4M 1.3M
2 Kimi Code Kimi K3 Max
29.1%
49.9%
$206 55h 44m 465.4M 2.8M
3 Codex GPT-5.6 Luna
29.1%
42.5%
$107 19h 52m 652.4M 2.2M
4 Codex GPT-5.6 Terra
27.3%
46.2%
$150 38h 15m 316.9M 1.8M
5 Claude Code Claude Fable 5
25.5%
42%
$1,826 25h 34m 1.1B 3.8M
6 Claude Code Seed 2.1 Pro -
24.1%
40.1%
$289 42h 44m 639.5M 3.2M
7 Claude Code Qwen3 8 XHigh
23.6%
46.4%
$149 57h 33m 392.1M 3.6M
8 Claude Code Kimi K3 Max
23.6%
44.9%
$144 64h 56m 297.8M 2.4M
9 Grok Build Grok 4.5 High
23.6%
44.6%
$114 29h 56m 230.1M 2.1M
10 Claude Code Claude Opus 4.8
23.6%
43.1%
$1,324 59h 37m 524.3M 8.2M
11 ALE-Claw GPT-5.5 High
23.6%
41.1%
$70 8h 28m 54.9M 734.8K
12 Codex GPT-5.5
22.4%
37.6%
$130 20h 57m 130.8M 1.2M
13 Cursor CLI Claude Opus 4.7 High
20%
37.6%
$842 14h 48m 215.6M 1.9M
14 Cursor CLI GPT-5.5 Medium
20%
34%
$50 17h 24m 36.4M 552.1K
15 OpenClaw GPT-5.4 High
19.4%
34.3%
$123 57h 48m 239.6M 2.9M
16 Claude Code GLM-5.2 Max
18.2%
37%
$390 35h 45m 338.8M 2.7M
17 ALE-Claw Claude Opus 4.7 High
18.2%
36.6%
$246 15h 31m 288.8M 1.7M
18 Droid GPT-5.5 High
18.2%
35.2%
$69 15h 37m 58.6M 792.2K
19 OpenClaw GPT-5.5 High
18.2%
33.4%
$101 22h 58m 98.6M 1.1M
20 Cursor CLI Composer 2.5 Adaptive
18.2%
32%
$68 25h 41m 130.0M 1.1M
21 OpenClaw Seed 2.1 Pro High
13.5%
23.6%
$171 46h 53m 572.9M 7.7M
22 Claude Code Claude Opus 4.7 High
12.7%
29.1%
$747 12h 5m 202.0M 1.5M
23 Gemini CLI Gemini 3.1 Pro High
12.7%
26.4%
$962 25h 34m 481.7M 1.7M
24 OpenClaw Qwen3 7-Max High
10.9%
28.5%
$280 35h 46m 585.1M 7.1M
25 OpenClaw Claude Opus 4.7 High
10.9%
27.5%
$370 29h 47m 163.0M 1.1M
26 OpenClaw DeepSeek V4 Pro High
10.9%
23.8%
$68 38h 12m 114.5M 1.5M
27 OpenClaw Gemini 3.1 Pro High
10.9%
23.6%
$1,392 62h 47m 1.6B 1.7M
28 ALE-Claw GPT-5.4 High
9.1%
22.9%
$179 22h 36m 595.7M 930.0K
29 OpenClaw GLM-5.1 High
9.1%
21.7%
$165 61h 57m 282.4M 2.2M
30 OpenClaw MiMo V2.5 High
9.1%
20.8%
$17 34h 51m 159.9M 1.4M
31 OpenClaw Qwen3 6-Plus High
8.2%
22.9%
$111 46h 37m 322.1M 3.1M
32 Grok CLI Grok 4.3 -
7.3%
18.3%
$95 10h 30m 74.0M 896.2K
33 OpenClaw Kimi K2.6 High
6.4%
18.2%
$41 73h 35m 93.6M 2.4M
34 OpenClaw Grok 4.3 High
3.6%
13.5%
$25 24h 45m 51.6M 1.0M
35 Droid Claude Opus 4.7 High
3.6%
10.9%
$167 3h 8m 39.9M 472.1K
36 OpenClaw MiniMax M2.7 High
3.6%
8.4%
$10 46h 17m 95.2M 1.7M
Rank Harness Model Effort Pass Rate Score Est. Cost Runtime Input Tokens Output Tokens
1 Claude Code Qwen3 8 XHigh
10.5%
24.1%
$217 95h 41m 559.3M 4.8M
2 Grok Build Grok 4.5 High
10.5%
23.7%
$115 47h 43m 223.7M 1.7M
3 Kimi Code Kimi K3 Max
10.5%
20.6%
$308 91h 53m 748.2M 3.1M
4 Claude Code Kimi K3 Max
7.9%
20.5%
$200 100h 52m 437.6M 2.6M
5 Claude Code Claude Fable 5
7.9%
16.5%
$2,091 30h 16m 1.6B 3.8M
6 Codex GPT-5.6 Sol
5.3%
19.4%
$325 38h 53m 352.3M 1.4M
7 Codex GPT-5.6 Luna
2.6%
15%
$197 31h 35m 1.4B 2.6M
8 Codex GPT-5.6 Terra
2.6%
14.9%
$264 50h 55m 678.1M 2.2M
9 Claude Code Claude Opus 4.8
2.6%
14.4%
$1,805 84h 44m 961.0M 7.3M
10 Codex GPT-5.5
2.6%
14.2%
$167 41h 36m 153.9M 905.0K
11 Claude Code GLM-5.2 Max
2.6%
12.9%
$638 47h 56m 782.7M 2.7M
12 ALE-Claw GPT-5.5 High
2.6%
12.8%
$107 13h 44m 126.3M 704.4K
13 Cursor CLI Claude Opus 4.7 High
2.6%
11.4%
$588 18h 31m 121.0M 1.2M
14 Droid GPT-5.5 High
2.6%
11.3%
$76 26h 27m 83.6M 551.8K
15 Cursor CLI GPT-5.5 Medium
2.6%
10.7%
$61 17h 12m 63.9M 456.4K
16 OpenClaw GPT-5.5 High
0%
10.9%
$155 24h 52m 178.2M 997.0K
17 Claude Code Seed 2.1 Pro -
0%
10.8%
$319 52h 3m 814.8M 2.8M
18 OpenClaw Seed 2.1 Pro High
0%
10.4%
$85 75h 30m 455.8M 4.8M
19 Claude Code Claude Opus 4.7 High
0%
10%
$685 17h 6m 170.1M 1.4M
20 Cursor CLI Composer 2.5 Adaptive
0%
8.8%
$78 34h 44m 151.0M 948.3K
21 ALE-Claw GPT-5.4 High
0%
8.1%
$100 17h 39m 321.4M 540.7K
22 Droid Claude Opus 4.7 High
0%
8%
$496 10h 18m 145.7M 920.1K
23 ALE-Claw Claude Opus 4.7 High
0%
7.9%
$594 43h 8m 707.3M 2.2M
24 OpenClaw GPT-5.4 High
0%
7.3%
$107 36h 16m 206.9M 2.5M
25 OpenClaw Qwen3 7-Max High
0%
6.4%
$173 33h 3m 377.8M 4.1M
26 OpenClaw GLM-5.1 High
0%
6.2%
$204 57h 10m 354.0M 2.5M
27 OpenClaw Qwen3 6-Plus High
0%
5%
$78 35h 46m 228.9M 1.9M
28 OpenClaw Claude Opus 4.7 High
0%
4.3%
$842 50h 0m 452.9M 1.5M
29 OpenClaw MiMo V2.5 High
0%
3.4%
$15 33h 3m 135.8M 1.3M
30 OpenClaw Gemini 3.1 Pro High
0%
3.1%
$1,200 45h 9m 1.4B 1.1M
31 OpenClaw DeepSeek V4 Pro High
0%
2.5%
$108 39h 7m 208.8M 1.7M
32 Grok CLI Grok 4.3 -
0%
2.3%
$71 11h 35m 56.2M 477.3K
33 OpenClaw Grok 4.3 High
0%
2.3%
$22 25h 4m 47.2M 621.3K
34 OpenClaw Kimi K2.6 High
0%
1.6%
$39 60h 46m 88.1M 1.8M
35 OpenClaw MiniMax M2.7 High
0%
1.3%
$9 44h 14m 87.7M 1.5M
36 Gemini CLI Gemini 3.1 Pro High
0%
0.9%
$733 40h 2m 497.3M 821.0K

*claude-fable-5: the variant Anthropic served during evaluation may differ from the published model's full capability tier, and re-runs cannot guarantee the higher-tier variant is selected. These numbers may understate the model's true ceiling. Learn more.

Pass rate vs input tokens

Efficiency frontier for this evaluation split. Up and to the left is better (higher pass rate at lower input-token spend). Several harness/model pairs reach the top of the split at a fraction of the cost of others; some pairs spend an order of magnitude more tokens without a corresponding pass-rate gain.

Hover any dot for the harness, model, and exact metrics. 

Industry coverage

Six representative task families across the 55 sub-industries ALE covers.
Image

Motion & VFX

Animation and visual effects production tasks in Adobe After Effects.
Image

3D modeling

3D model creation and editing tasks in Siemens NX.
Image

Game development

Scene setup, asset placement, and rendering tasks in Unreal Engine.
Image

Mold flow analysis

Simulation and mold flow analysis tasks in Moldex3D manufacturing software.
Image

Architectural modeling

3D modeling and energy analysis workflows in Rhino 3D for urban design.
Image

Brain imaging

Neuroimaging analysis and brain structure segmentation tasks in FSLeyes.

Sample tasks

A selection from the 147 public ALE-V1 tasks across 14 task categories. Each task ships with a sandboxed environment, a hidden reference, and a deterministic grader. Slugs link to the task source.

business finance

sec_10k_financial_parsing

Parse a SEC 10-K filing into a structured financial schema. Multi-step extraction, table normalization, and cross-reference validation against the original document.

business finance

financial_stmt_reconstruction_aapl_fy2024
Reconstruct Apple’s FY2024 financial statement from primary disclosure documents. Validates whether the agent surfaces the exact reported figures and footnote-relevant adjustments.
engineering
mold-flow / 220089
Set up a Moldex3D mold-flow simulation, run it to convergence, and report fill time / pressure metrics matching the held-out reference run.

health medicine

Clinical_Variant_Annotation

Annotate a clinical variant set using standard pipelines (VEP, ClinVar, etc.) and produce a report graded against a curated reference.

life sciences

WGS_Variant_Calling
Run a whole-genome sequencing variant-calling pipeline and produce VCF output. Scored on precision and recall against a held-out truth VCF.

computing math

k8s_payment_api_root_cause_analysis
Diagnose a failing payment API in a Kubernetes cluster. Multi-hop investigation across logs, metrics, manifests, and traces, scored on the correct root-cause identification.

visual media

video_storyboard_001
Build a shot-by-shot video storyboard from a brief, formatted to industry conventions. Graded on coverage, continuity, and adherence to the reference shot list.
legal
legal_dr_fees_01
Compute legal fees from a billing register according to jurisdictional rules. Tests structured extraction plus rule-following against an authoritative reference total.

Methodology

Metrics

Pass Rate — fraction of tasks the agent fully completed (strict success). Score — average graded outcome across all tasks, including partial credit. Both computed by deterministic graders against hidden references.

Verifiable Outcomes
Hidden references plus deterministic graders, not LLM-as-a-judge. Tasks sourced from real professional workflows (After Effects, Siemens NX, Unreal Engine, Moldex3D, Rhino 3D, FSLeyes, and 49 more applications) and validated by domain experts before inclusion.
Rolling Evaluation

Every ~6 months, a new public subset releases with fresh instances. Private tasks rotate into the public pool, retired public tasks rotate out, and held-out private tasks score the official leaderboard, to limit benchmark leakage.

Reference Harnesses

Two open harnesses ship with the framework: the official Claude Code CLI and the in-tree OpenClaw harness. Submissions also include Codex, Cursor CLI, Droid, Gemini CLI, Grok CLI, and the ALE Claw reference harness.

Acknowledgments

Agents’ Last Exam is co-led by UC Berkeley RDI and the RDI Foundation, with funding support and contributions from Snorkel AI via the Open Benchmarks Grants program. The benchmark draws task contributions from 300+ industry experts across 44 academic institutions (MIT, Harvard, Stanford, UC Berkeley, Oxford, CMU, Caltech, ETH Zurich, Yale, Columbia, and more) and industry organizations including Goldman Sachs, JPMorgan, Morgan Stanley, PIMCO, Meta, Amazon, Adobe, Oracle, Hippocratic AI, and HubSpot.

Advisory Committee includes George Em Karniadakis (Brown), Tapio Schneider (Caltech), Teresa Head-Gordon (UC Berkeley), Laure Zanna (NYU), Jack Gallant (UC Berkeley), Tarek Zohdi (UC Berkeley), Ida Sim (UCSF), Arvind Rao (U Michigan), Kaan Ozbay (NYU), Carl Boettiger (UC Berkeley), Kyle Steinfeld (UC Berkeley), Yamini Rangan (HubSpot), and Bradley Rothenberg (nTop).

FAQs

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

Open Benchmarks Grants

Terminal-Bench 3.0

The next frontier benchmark for agent work. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.

By Resolution Rate
1
Image
Opus 5 (max) • mini-SWE-agent
43.5%
2
Image
GPT-5.6 Sol (max) • Codex
32.7%
3
Image
Fable 5 (max) • Claude Code
31.7%

Senior SWE-bench

A benchmark for evaluating coding agents on senior-level engineering work: building features from realistic instructions, investigating bugs that require runtime investigation, and shipping code that aligns to existing codebase conventions.

Tasteful Solve Rate
1
Image
Claude Fable 5
34.7%
2
Image
Claude Opus 5
34.7%
3
Image
GPT-5.6 Sol
34.7%
Open Benchmarks Grants

OSWorld 2.0

A benchmark for evaluating computer-use agents on long-horizon, real-world workflows: 108 authentic tasks across 31 self-hosted web environments and professional desktop applications.

By Binary Accuracy (300 steps)
1
Image
GPT-5.5 · xhigh
13%
2
Image
Claude Opus 4.7 · max
13%
3
Image
Claude Sonnet 4.6 · medium
8.3%
Open Benchmarks Grants

SlopCode Bench

Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.

Top Models by Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3-Codex
26.02%
3
Image
GPT-5.4
23.47%
Open Benchmarks Grants

Continual Learning Bench

Evaluates whether AI systems improve from prior experience across sequential, stateful tasks, measuring real in-context learning, not just raw capability.

Top Systems (Agg. Reward)
1
Image
ICL · Claude Sonnet 4.6
+0.196
2
Image
ICL · GPT-5.4
+0.189
3
Image
Claude Code · Sonnet 4.6
+0.185
Open Benchmarks Grants

Terminal-Bench 2.1

Terminal agent evaluation led by Stanford University and Laude Institute. v2.1 fixes 28 tasks from 2.0 and introduces continuous validation.

Top Submissions
1
Image
Claude Code · Claude 5 Fable
83.8%
2
Image
Codex CLI · GPT-5.5
83.1%
3
Image
Terminus 2 · Claude 5 Fable
80.4%
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.