Terminal-Bench-Science
Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work. Scientists, not model developers or vendors, set the bar for scientific capability in AI.
Overview
Terminal-Bench-Science evaluates agents on real scientific research workflows rather than software repository tasks. Version 0.1, the benchmark’s first public release, covers 70 tasks across five domain areas, each sourced from work a practicing researcher actually performed.
Headline finding: the strongest model evaluated, Claude Opus 5, resolves 30.0% of tasks in this release, well below the 50 to 80% range frontier models reach on general coding benchmarks. In this evaluation, authentic scientific workflows remain considerably harder for agents than general software engineering.
Snorkel supports Terminal-Bench Science through the Open Benchmarks Grants program.
At a glance
70
tasks in the v0.1 release
5
9
30.0%
Leaderboard
| Rank | Model | Agent | Effort | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Claude Code | max |
30%
±3.2
|
2026-07-24 | 7.3B | $6.99K |
| 2 | GPT-5.6 Sol | Codex | max |
22.4%
±2.9
|
2026-07-09 | 8.4B | $4.22K |
| 3 | Claude Fable 5 | Claude Code | max |
21.4%
±2.8
|
2026-06-09 | 6.4B | $14.18K |
| 4 | Claude Opus 4.8 | Claude Code | max |
10.5%
±2.1
|
2026-05-28 | 6.8B | $5.84K |
| 5 | GPT-5.6 Terra | Codex | max |
8.6%
±1.9
|
2026-07-09 | 7.6B | $1.98K |
| 6 | GLM 5.3 | Claude Code | max |
8.1%
±1.9
|
2026-08-14 | 8.5B | $2.73K |
| 7 | Kimi K3 | Claude Code | max |
7.1%
±1.8
|
2026-07-16 | 3.2B | $1.52K |
| 8 | Grok 4.6 | Grok Build | xhigh |
7.1%
±1.8
|
2026-08-12 | 3.7B | $3.34K |
| 9 | GPT-5.6 Luna | Codex | max |
3.3%
±1.2
|
2026-07-09 | 14.2B | $383.20 |
| Rank | Model | Agent | Effort | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Claude Code | max |
29.8%
|
2026-07-24 | 7.3B | $6.99K |
| 2 | GPT-5.6 Sol | Codex | max |
17.5%
|
2026-07-09 | 8.4B | $4.22K |
| 3 | Claude Fable 5 | Claude Code | max |
17.5%
|
2026-06-09 | 6.4B | $14.18K |
| 4 | GLM 5.3 | Claude Code | max |
10.5%
|
2026-08-14 | 8.5B | $2.73K |
| 5 | Claude Opus 4.8 | Claude Code | max |
7%
|
2026-05-28 | 6.8B | $5.84K |
| 6 | GPT-5.6 Terra | Codex | max |
7%
|
2026-07-09 | 7.6B | $1.98K |
| 7 | Kimi K3 | Claude Code | max |
7%
|
2026-07-16 | 3.2B | $1.52K |
| 8 | Grok 4.6 | Grok Build | xhigh |
7%
|
2026-08-12 | 3.7B | $3.34K |
| 9 | GPT-5.6 Luna | Codex | max |
1.8%
|
2026-07-09 | 14.2B | $383.20 |
| Rank | Model | Agent | Effort | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Claude Code | max |
45.8%
|
2026-07-24 | 7.3B | $6.99K |
| 2 | Claude Fable 5 | Claude Code | max |
25%
|
2026-06-09 | 6.4B | $14.18K |
| 3 | GPT-5.6 Sol | Codex | max |
20.8%
|
2026-07-09 | 8.4B | $4.22K |
| 4 | Kimi K3 | Claude Code | max |
12.5%
|
2026-07-16 | 3.2B | $1.52K |
| 5 | GPT-5.6 Terra | Codex | max |
8.3%
|
2026-07-09 | 7.6B | $1.98K |
| 6 | Claude Opus 4.8 | Claude Code | max |
4.2%
|
2026-05-28 | 6.8B | $5.84K |
| 7 | GLM 5.3 | Claude Code | max |
4.2%
|
2026-08-14 | 8.5B | $2.73K |
| 8 | Grok 4.6 | Grok Build | xhigh |
4.2%
|
2026-08-12 | 3.7B | $3.34K |
| 9 | GPT-5.6 Luna | Codex | max |
0%
|
2026-07-09 | 14.2B | $383.20 |
| Rank | Model | Agent | Effort | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Claude Code | max |
29.6%
|
2026-07-24 | 7.3B | $6.99K |
| 2 | GPT-5.6 Sol | Codex | max |
14.8%
|
2026-07-09 | 8.4B | $4.22K |
| 3 | Grok 4.6 | Grok Build | xhigh |
14.8%
|
2026-08-12 | 3.7B | $3.34K |
| 4 | Claude Fable 5 | Claude Code | max |
11.1%
|
2026-06-09 | 6.4B | $14.18K |
| 5 | Claude Opus 4.8 | Claude Code | max |
7.4%
|
2026-05-28 | 6.8B | $5.84K |
| 6 | GPT-5.6 Terra | Codex | max |
7.4%
|
2026-07-09 | 7.6B | $1.98K |
| 7 | GLM 5.3 | Claude Code | max |
7.4%
|
2026-08-14 | 8.5B | $2.73K |
| 8 | GPT-5.6 Luna | Codex | max |
7.4%
|
2026-07-09 | 14.2B | $383.20 |
| 9 | Kimi K3 | Claude Code | max |
3.7%
|
2026-07-16 | 3.2B | $1.52K |
| Rank | Model | Agent | Effort | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Claude Code | max |
33.3%
|
2026-06-09 | 6.4B | $14.18K |
| 2 | GPT-5.6 Sol | Codex | max |
31.4%
|
2026-07-09 | 8.4B | $4.22K |
| 3 | Claude Opus 5 | Claude Code | max |
25.5%
|
2026-07-24 | 7.3B | $6.99K |
| 4 | GPT-5.6 Terra | Codex | max |
9.8%
|
2026-07-09 | 7.6B | $1.98K |
| 5 | Claude Opus 4.8 | Claude Code | max |
7.8%
|
2026-05-28 | 6.8B | $5.84K |
| 6 | GPT-5.6 Luna | Codex | max |
7.8%
|
2026-07-09 | 14.2B | $383.20 |
| 7 | GLM 5.3 | Claude Code | max |
5.9%
|
2026-08-14 | 8.5B | $2.73K |
| 8 | Grok 4.6 | Grok Build | xhigh |
5.9%
|
2026-08-12 | 3.7B | $3.34K |
| 9 | Kimi K3 | Claude Code | max |
2%
|
2026-07-16 | 3.2B | $1.52K |
| Rank | Model | Agent | Effort | Resolution Rate | Release Date | Tokens | Cost |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Claude Code | max |
27.5%
|
2026-07-24 | 7.3B | $6.99K |
| 2 | GPT-5.6 Sol | Codex | max |
23.5%
|
2026-07-09 | 8.4B | $4.22K |
| 3 | Claude Opus 4.8 | Claude Code | max |
21.6%
|
2026-05-28 | 6.8B | $5.84K |
| 4 | Claude Fable 5 | Claude Code | max |
17.6%
|
2026-06-09 | 6.4B | $14.18K |
| 5 | Kimi K3 | Claude Code | max |
11.8%
|
2026-07-16 | 3.2B | $1.52K |
| 6 | GPT-5.6 Terra | Codex | max |
9.8%
|
2026-07-09 | 7.6B | $1.98K |
| 7 | GLM 5.3 | Claude Code | max |
9.8%
|
2026-08-14 | 8.5B | $2.73K |
| 8 | Grok 4.6 | Grok Build | xhigh |
5.9%
|
2026-08-12 | 3.7B | $3.34K |
| 9 | GPT-5.6 Luna | Codex | max |
0%
|
2026-07-09 | 14.2B | $383.20 |
Terminal-Bench-Science 0.1 Pareto Frontier
TAXONOMY
Domain areas
The v0.1 release spans five domain areas, unevenly weighted toward life and physical sciences. Resolution rates above are aggregated across all 70 tasks.
19
Life sciences
17
8
17
9
Sample tasks
A sample of task titles from the v0.1 release. Browse the full contribution pipeline on the task dashboard →
CHEMISTRY
Constraint-Consistent Internal-Coordinate Embedding for RDKit
Implement constrained RDKit conformer generation under distance, angle, signed-torsion, and ambiguous-distance restraints.
ASTRONOMY
BIOLOGY
Genomic model ranking
MEDICINE
APPLIED MATHEMATICS
FORMAL MATHEMATICS
GEOSCIENCES
MECHANICAL ENGINEERING
How tasks are graded
Every task in Terminal-Bench-Science goes through a review pipeline before it is eligible to be merged into a release, and every evaluation run separates the agent’s working environment from the environment that grades it.
Resources
Acknowledgments
Terminal-Bench-Science is a collaboration between Stanford University, the Laude Institute, and domain experts across scientific disciplines and research institutions worldwide, led by Steven Dillmann (Stanford University), Sanmi Koyejo (Stanford University), and Ludwig Schmidt (Stanford University, Anthropic)
Snorkel supports Terminal-Bench Science through the Open Benchmarks Grants program.

