Terminal-Bench-Science
At a glance
70
5
68.1%
top resolution rate
(GPT-6 Astra)
Terminal-Bench-Science 0.1 Pareto Frontier
- GPT-6 Astra
- Opus 5.5
- Fable 5.1
- Opus 5
- GPT-5.6 Sol
Leaderboard
| Rank | Model | Agent | Effort | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra | Codex | max |
68.1%
±3.2
|
2026-09-03 | 2.3B | $4,997 |
| 2 | Opus 5.5 | Claude Code | max |
63.3%
±3.3
|
2026-09-22 | 8.7B | $4,874 |
| 3 | Fable 5.1 | Claude Code | max |
40%
±3.4
|
2026-09-01 | 3.8B | $6.25K |
| 4 | Opus 5 | Claude Code | max |
30%
±3.2
|
2026-07-24 | 7.3B | $6.99K |
| 5 | GPT-5.6 Sol | Codex | max |
22.4%
±2.9
|
2026-07-09 | 8.4B | $4.22K |
| 6 | Fable 5 | Claude Code | max |
21.4%
±2.8
|
2026-06-09 | 6.4B | $14.18K |
| 7 | DeepSeek V4.1 Flash | Codex | max |
15.7%
±2.5
|
2026-09-10 | 14.9B | $385.93 |
| 8 | Grok 4.7 | Grok Build | xhigh |
14.3%
±2.4
|
2026-09-21 | 4.2B | $3,071 |
| 9 | Gemini 3.8 Flash | mini-SWE-agent | high |
12.4%
±2.3
|
2026-09-02 | 10.5B | $1.12K |
| 10 | Opus 4.8 | Claude Code | max |
10.5%
±2.1
|
2026-05-28 | 6.8B | $5.84K |
| 11 | GPT-5.6 Terra | Codex | max |
8.6%
±1.9
|
2026-07-09 | 7.6B | $1.98K |
| 12 | GLM 5.3 | Claude Code | max |
8.1%
±1.9
|
2026-08-14 | 8.5B | $2.73K |
| 13 | Kimi K3 | Claude Code | max |
7.1%
±1.8
|
2026-07-16 | 3.2B | $1.52K |
| 14 | Grok 4.6 | Grok Build | xhigh |
7.1%
±1.8
|
2026-08-12 | 3.7B | $3.34K |
| 15 | Gemini 3.7 Flash | mini-SWE-agent | high |
5.7%
±1.6
|
2026-08-13 | 12.0B | $1.30K |
| 16 | DeepSeek V4 Pro | Codex | max |
3.8%
±1.3
|
2026-08-13 | 11.2B | $1,120.15 |
| 17 | GPT-5.6 Luna | Codex | max |
3.3%
±1.2
|
2026-07-09 | 14.2B | $383.20 |
| Rank | Model | Agent | Effort | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra | Codex | max |
59.6%
±6.5
|
2026-09-03 | 499.9M | $1.03K |
| 2 | Opus 5.5 | Claude Code | max |
54.4%
±6.6
|
2026-09-22 | 1.9B | $970.39 |
| 3 | Fable 5.1 | Claude Code | max |
38.6%
±6.4
|
2026-09-01 | 751.4M | $1.11K |
| 4 | Opus 5 | Claude Code | max |
29.8%
±6.1
|
2026-07-24 | 1.3B | $1.32K |
| 5 | GPT-5.6 Sol | Codex | max |
17.5%
±5
|
2026-07-09 | 911.2M | $513.19 |
| 6 | Fable 5 | Claude Code | max |
17.5%
±5
|
2026-06-09 | 864.3M | $1.96K |
| 7 | Gemini 3.8 Flash | mini-SWE-agent | high |
14%
±4.6
|
2026-09-02 | 947.5M | $129.49 |
| 8 | Grok 4.7 | Grok Build | xhigh |
12.3%
±4.3
|
2026-09-21 | 516.7M | $377.22 |
| 9 | GLM 5.3 | Claude Code | max |
10.5%
±4.1
|
2026-08-14 | 1.5B | $1.31K |
| 10 | DeepSeek V4.1 Flash | Codex | max |
10.5%
±4.1
|
2026-09-10 | 2.5B | $64.23 |
| 11 | Gemini 3.7 Flash | mini-SWE-agent | high |
8.8%
±3.7
|
2026-08-13 | 704.8M | $100.54 |
| 12 | Grok 4.6 | Grok Build | xhigh |
7%
±3.4
|
2026-08-12 | 367.5M | $281.65 |
| 13 | GPT-5.6 Terra | Codex | max |
7%
±3.4
|
2026-07-09 | 1B | $297.96 |
| 14 | Kimi K3 | Claude Code | max |
7%
±3.4
|
2026-07-16 | 612.8M | $591.34 |
| 15 | Opus 4.8 | Claude Code | max |
7%
±3.4
|
2026-05-28 | 780.5M | $806.35 |
| 16 | GPT-5.6 Luna | Codex | max |
1.8%
±1.7
|
2026-07-09 | 2.9B | $83.51 |
| 17 | DeepSeek V4 Pro | Codex | max |
0%
|
2026-08-13 | 2.1B | $107.43 |
| Rank | Model | Agent | Effort | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|---|
| 1 | Opus 5.5 | Claude Code | max |
60.8%
±6.8
|
2026-09-22 | 2.1B | $1.08K |
| 2 | GPT-6 Astra | Codex | max |
58.8%
±6.9
|
2026-09-03 | 558.1M | $1.22K |
| 3 | Fable 5.1 | Claude Code | max |
37.3%
±6.8
|
2026-09-01 | 1B | $1.82K |
| 4 | Opus 5 | Claude Code | max |
27.5%
±6.2
|
2026-07-24 | 2.1B | $1.85K |
| 5 | DeepSeek V4.1 Flash | Codex | max |
27.5%
±6.2
|
2026-09-10 | 4.6B | $118.82 |
| 6 | Gemini 3.8 Flash | mini-SWE-agent | high |
23.5%
±5.9
|
2026-09-02 | 2.8B | $288.67 |
| 7 | GPT-5.6 Sol | Codex | max |
23.5%
±5.9
|
2026-07-09 | 2.4B | $1.14K |
| 8 | Opus 4.8 | Claude Code | max |
21.6%
±5.8
|
2026-05-28 | 2B | $1.6K |
| 9 | Fable 5 | Claude Code | max |
17.6%
±5.3
|
2026-06-09 | 1.5B | $3.58K |
| 10 | Grok 4.7 | Grok Build | xhigh |
15.7%
±5.1
|
2026-09-21 | 913.1M | $680.14 |
| 11 | Kimi K3 | Claude Code | max |
11.8%
±4.5
|
2026-07-16 | 799.3M | $536.23 |
| 12 | GPT-5.6 Terra | Codex | max |
9.8%
±4.2
|
2026-07-09 | 3.3B | $791.88 |
| 13 | GLM 5.3 | Claude Code | max |
9.8%
±4.2
|
2026-08-14 | 2.3B | $1.73K |
| 14 | Gemini 3.7 Flash | mini-SWE-agent | high |
7.8%
±3.8
|
2026-08-13 | 4.7B | $490.74 |
| 15 | Grok 4.6 | Grok Build | xhigh |
5.9%
±3.3
|
2026-08-12 | 646.9M | $539.59 |
| 16 | DeepSeek V4 Pro | Codex | max |
3.9%
±2.7
|
2026-08-13 | 2.8B | $140.45 |
| 17 | GPT-5.6 Luna | Codex | max |
0%
|
2026-07-09 | 4.8B | $116.84 |
| Rank | Model | Agent | Effort | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra | Codex | max |
79.2%
±8.3
|
2026-09-03 | 293.3M | $595.78 |
| 2 | Opus 5.5 | Claude Code | max |
62.5%
±9.9
|
2026-09-22 | 1.3B | $638.23 |
| 3 | Opus 5 | Claude Code | max |
45.8%
±10.2
|
2026-07-24 | 897.5M | $1.15K |
| 4 | Fable 5.1 | Claude Code | max |
37.5%
±9.9
|
2026-09-01 | 431.2M | $746.48 |
| 5 | Grok 4.7 | Grok Build | xhigh |
25%
±8.8
|
2026-09-21 | 437.3M | $314.11 |
| 6 | Fable 5 | Claude Code | max |
25%
±8.8
|
2026-06-09 | 602.6M | $1.62K |
| 7 | GPT-5.6 Sol | Codex | max |
20.8%
±8.3
|
2026-07-09 | 1.2B | $584.56 |
| 8 | Gemini 3.8 Flash | mini-SWE-agent | high |
12.5%
±6.8
|
2026-09-02 | 1.1B | $117.58 |
| 9 | Kimi K3 | Claude Code | max |
12.5%
±6.8
|
2026-07-16 | 380.1M | $342.60 |
| 10 | GPT-5.6 Terra | Codex | max |
8.3%
±5.6
|
2026-07-09 | 1.2B | $302.41 |
| 11 | DeepSeek V4.1 Flash | Codex | max |
8.3%
±5.6
|
2026-09-10 | 1.8B | $39.17 |
| 12 | Grok 4.6 | Grok Build | xhigh |
4.2%
±4.1
|
2026-08-12 | 244.9M | $217.33 |
| 13 | Opus 4.8 | Claude Code | max |
4.2%
±4.1
|
2026-05-28 | 753.7M | $694.26 |
| 14 | GLM 5.3 | Claude Code | max |
4.2%
±4.1
|
2026-08-14 | 1.6B | $1.4K |
| 15 | GPT-5.6 Luna | Codex | max |
0%
|
2026-07-09 | 1.6B | $41.41 |
| 16 | Gemini 3.7 Flash | mini-SWE-agent | high |
0%
|
2026-08-13 | 503.6M | $63.51 |
| 17 | DeepSeek V4 Pro | Codex | max |
0%
|
2026-08-13 | 1.5B | $71.82 |
| Rank | Model | Agent | Effort | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra | Codex | max |
70.4%
±8.8
|
2026-09-03 | 275.5M | $698.22 |
| 2 | Opus 5.5 | Claude Code | max |
63%
±9.3
|
2026-09-22 | 1.7B | $1.03K |
| 3 | Fable 5.1 | Claude Code | max |
48.1%
±9.6
|
2026-09-01 | 633.4M | $976.65 |
| 4 | Opus 5 | Claude Code | max |
29.6%
±8.8
|
2026-07-24 | 1.2B | $1.04K |
| 5 | Grok 4.7 | Grok Build | xhigh |
25.9%
±8.4
|
2026-09-21 | 464.6M | $343.58 |
| 6 | DeepSeek V4.1 Flash | Codex | max |
18.5%
±7.5
|
2026-09-10 | 3B | $69.75 |
| 7 | Grok 4.6 | Grok Build | xhigh |
14.8%
±6.8
|
2026-08-12 | 276.7M | $231.57 |
| 8 | GPT-5.6 Sol | Codex | max |
14.8%
±6.8
|
2026-07-09 | 569.2M | $315.06 |
| 9 | Gemini 3.8 Flash | mini-SWE-agent | high |
11.1%
±6
|
2026-09-02 | 520.2M | $70.26 |
| 10 | DeepSeek V4 Pro | Codex | max |
11.1%
±6
|
2026-08-13 | 2.2B | $108.40 |
| 11 | Fable 5 | Claude Code | max |
11.1%
±6
|
2026-06-09 | 1.3B | $2.39K |
| 12 | GPT-5.6 Luna | Codex | max |
7.4%
±5
|
2026-07-09 | 1.2B | $35.82 |
| 13 | Gemini 3.7 Flash | mini-SWE-agent | high |
7.4%
±5
|
2026-08-13 | 377.2M | $56.41 |
| 14 | GPT-5.6 Terra | Codex | max |
7.4%
±5
|
2026-07-09 | 464.1M | $141.00 |
| 15 | GLM 5.3 | Claude Code | max |
7.4%
±5
|
2026-08-14 | 992.1M | $788.27 |
| 16 | Opus 4.8 | Claude Code | max |
7.4%
±5
|
2026-05-28 | 1.1B | $831.49 |
| 17 | Kimi K3 | Claude Code | max |
3.7%
±3.6
|
2026-07-16 | 536.3M | $354.73 |
| Rank | Model | Agent | Effort | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra | Codex | max |
80.4%
±5.6
|
2026-09-03 | 661.6M | $1.45K |
| 2 | Opus 5.5 | Claude Code | max |
76.5%
±5.9
|
2026-09-22 | 1.7B | $1.16K |
| 3 | Fable 5.1 | Claude Code | max |
41.2%
±6.9
|
2026-09-01 | 891.4M | $1.6K |
| 4 | Fable 5 | Claude Code | max |
33.3%
±6.6
|
2026-06-09 | 2.1B | $4.63K |
| 5 | GPT-5.6 Sol | Codex | max |
31.4%
±6.5
|
2026-07-09 | 3.3B | $1.66K |
| 6 | Opus 5 | Claude Code | max |
25.5%
±6.1
|
2026-07-24 | 1.7B | $1.63K |
| 7 | DeepSeek V4.1 Flash | Codex | max |
11.8%
±4.5
|
2026-09-10 | 2.9B | $93.97 |
| 8 | GPT-5.6 Terra | Codex | max |
9.8%
±4.2
|
2026-07-09 | 1.6B | $450.66 |
| 9 | GPT-5.6 Luna | Codex | max |
7.8%
±3.8
|
2026-07-09 | 3.7B | $105.62 |
| 10 | Opus 4.8 | Claude Code | max |
7.8%
±3.8
|
2026-05-28 | 2.2B | $1.9K |
| 11 | DeepSeek V4 Pro | Codex | max |
5.9%
±3.3
|
2026-08-13 | 2.7B | $146.13 |
| 12 | GLM 5.3 | Claude Code | max |
5.9%
±3.3
|
2026-08-14 | 2.1B | $1.57K |
| 13 | Grok 4.6 | Grok Build | xhigh |
5.9%
±3.3
|
2026-08-12 | 2.2B | $2.07K |
| 14 | Grok 4.7 | Grok Build | xhigh |
3.9%
±2.7
|
2026-09-21 | 1.9B | $1.36K |
| 15 | Gemini 3.7 Flash | mini-SWE-agent | high |
2%
±1.9
|
2026-08-13 | 5.7B | $585.39 |
| 16 | Kimi K3 | Claude Code | max |
2%
±1.9
|
2026-07-16 | 863.2M | $714.76 |
| 17 | Gemini 3.8 Flash | mini-SWE-agent | high |
0%
|
2026-09-02 | 5.2B | $510.98 |
TAXONOMY
Domain areas
The v0.1 release spans five domain areas, unevenly weighted toward life and physical sciences. Resolution rates above are aggregated across all 70 tasks.
19
Life sciences
17
8
17
9
Sample tasks
A sample of task titles from the v0.1 release. Browse the full contribution pipeline on the task dashboard →
CHEMISTRY
Constraint-Consistent Internal-Coordinate Embedding for RDKit
Implement constrained RDKit conformer generation under distance, angle, signed-torsion, and ambiguous-distance restraints.
ASTRONOMY
BIOLOGY
Genomic model ranking
MEDICINE
APPLIED MATHEMATICS
FORMAL MATHEMATICS
GEOSCIENCES
MECHANICAL ENGINEERING
How tasks are graded
Every task in Terminal-Bench-Science goes through a review pipeline before it is eligible to be merged into a release, and every evaluation run separates the agent’s working environment from the environment that grades it.
Resources
Behind the benchmark
Behind the benchmark
Terminal-Bench-Science evaluates agents on real scientific research workflows rather than software repository tasks. Version 0.1, the benchmark's first public release, covers 70 tasks across five domain areas, each sourced from work a practicing researcher actually performed.
Acknowledgments
Terminal-Bench-Science is a collaboration between Stanford University, the Laude Institute, and domain experts across scientific disciplines and research institutions worldwide, led by Steven Dillmann (Stanford University), Sanmi Koyejo (Stanford University), and Ludwig Schmidt (Stanford University, Anthropic)
Snorkel supports Terminal-Bench Science through the Open Benchmarks Grants program.

