Science & Research
Open Benchmarks Grants

Terminal-Bench-Science

Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work. Scientists, not model developers or vendors, set the bar for scientific capability in AI.
Built with
Image
Image
Image
Image

At a glance

70

tasks in the v0.1 release

5

domain areas covered

68.1% 

top resolution rate
(GPT-6 Astra)

Terminal-Bench-Science 0.1 Pareto Frontier

  • GPT-6 Astra
  • Opus 5.5
  • Fable 5.1
  • Opus 5
  • GPT-5.6 Sol
GPT-6 Astra ×
Opus 5.5 ×
Fable 5.1 ×
Opus 5 ×
GPT-5.6 Sol ×

Leaderboard

Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 GPT-6 Astra Codex max
68.1% ±3.2
2026-09-03 2.3B $4,997
2 Opus 5.5 Claude Code max
63.3% ±3.3
2026-09-22 8.7B $4,874
3 Fable 5.1 Claude Code max
40% ±3.4
2026-09-01 3.8B $6.25K
4 Opus 5 Claude Code max
30% ±3.2
2026-07-24 7.3B $6.99K
5 GPT-5.6 Sol Codex max
22.4% ±2.9
2026-07-09 8.4B $4.22K
6 Fable 5 Claude Code max
21.4% ±2.8
2026-06-09 6.4B $14.18K
7 DeepSeek V4.1 Flash Codex max
15.7% ±2.5
2026-09-10 14.9B $385.93
8 Grok 4.7 Grok Build xhigh
14.3% ±2.4
2026-09-21 4.2B $3,071
9 Gemini 3.8 Flash mini-SWE-agent high
12.4% ±2.3
2026-09-02 10.5B $1.12K
10 Opus 4.8 Claude Code max
10.5% ±2.1
2026-05-28 6.8B $5.84K
11 GPT-5.6 Terra Codex max
8.6% ±1.9
2026-07-09 7.6B $1.98K
12 GLM 5.3 Claude Code max
8.1% ±1.9
2026-08-14 8.5B $2.73K
13 Kimi K3 Claude Code max
7.1% ±1.8
2026-07-16 3.2B $1.52K
14 Grok 4.6 Grok Build xhigh
7.1% ±1.8
2026-08-12 3.7B $3.34K
15 Gemini 3.7 Flash mini-SWE-agent high
5.7% ±1.6
2026-08-13 12.0B $1.30K
16 DeepSeek V4 Pro Codex max
3.8% ±1.3
2026-08-13 11.2B $1,120.15
17 GPT-5.6 Luna Codex max
3.3% ±1.2
2026-07-09 14.2B $383.20
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 GPT-6 Astra Codex max
59.6% ±6.5
2026-09-03 499.9M $1.03K
2 Opus 5.5 Claude Code max
54.4% ±6.6
2026-09-22 1.9B $970.39
3 Fable 5.1 Claude Code max
38.6% ±6.4
2026-09-01 751.4M $1.11K
4 Opus 5 Claude Code max
29.8% ±6.1
2026-07-24 1.3B $1.32K
5 GPT-5.6 Sol Codex max
17.5% ±5
2026-07-09 911.2M $513.19
6 Fable 5 Claude Code max
17.5% ±5
2026-06-09 864.3M $1.96K
7 Gemini 3.8 Flash mini-SWE-agent high
14% ±4.6
2026-09-02 947.5M $129.49
8 Grok 4.7 Grok Build xhigh
12.3% ±4.3
2026-09-21 516.7M $377.22
9 GLM 5.3 Claude Code max
10.5% ±4.1
2026-08-14 1.5B $1.31K
10 DeepSeek V4.1 Flash Codex max
10.5% ±4.1
2026-09-10 2.5B $64.23
11 Gemini 3.7 Flash mini-SWE-agent high
8.8% ±3.7
2026-08-13 704.8M $100.54
12 Grok 4.6 Grok Build xhigh
7% ±3.4
2026-08-12 367.5M $281.65
13 GPT-5.6 Terra Codex max
7% ±3.4
2026-07-09 1B $297.96
14 Kimi K3 Claude Code max
7% ±3.4
2026-07-16 612.8M $591.34
15 Opus 4.8 Claude Code max
7% ±3.4
2026-05-28 780.5M $806.35
16 GPT-5.6 Luna Codex max
1.8% ±1.7
2026-07-09 2.9B $83.51
17 DeepSeek V4 Pro Codex max
0%
2026-08-13 2.1B $107.43
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Opus 5.5 Claude Code max
60.8% ±6.8
2026-09-22 2.1B $1.08K
2 GPT-6 Astra Codex max
58.8% ±6.9
2026-09-03 558.1M $1.22K
3 Fable 5.1 Claude Code max
37.3% ±6.8
2026-09-01 1B $1.82K
4 Opus 5 Claude Code max
27.5% ±6.2
2026-07-24 2.1B $1.85K
5 DeepSeek V4.1 Flash Codex max
27.5% ±6.2
2026-09-10 4.6B $118.82
6 Gemini 3.8 Flash mini-SWE-agent high
23.5% ±5.9
2026-09-02 2.8B $288.67
7 GPT-5.6 Sol Codex max
23.5% ±5.9
2026-07-09 2.4B $1.14K
8 Opus 4.8 Claude Code max
21.6% ±5.8
2026-05-28 2B $1.6K
9 Fable 5 Claude Code max
17.6% ±5.3
2026-06-09 1.5B $3.58K
10 Grok 4.7 Grok Build xhigh
15.7% ±5.1
2026-09-21 913.1M $680.14
11 Kimi K3 Claude Code max
11.8% ±4.5
2026-07-16 799.3M $536.23
12 GPT-5.6 Terra Codex max
9.8% ±4.2
2026-07-09 3.3B $791.88
13 GLM 5.3 Claude Code max
9.8% ±4.2
2026-08-14 2.3B $1.73K
14 Gemini 3.7 Flash mini-SWE-agent high
7.8% ±3.8
2026-08-13 4.7B $490.74
15 Grok 4.6 Grok Build xhigh
5.9% ±3.3
2026-08-12 646.9M $539.59
16 DeepSeek V4 Pro Codex max
3.9% ±2.7
2026-08-13 2.8B $140.45
17 GPT-5.6 Luna Codex max
0%
2026-07-09 4.8B $116.84
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 GPT-6 Astra Codex max
79.2% ±8.3
2026-09-03 293.3M $595.78
2 Opus 5.5 Claude Code max
62.5% ±9.9
2026-09-22 1.3B $638.23
3 Opus 5 Claude Code max
45.8% ±10.2
2026-07-24 897.5M $1.15K
4 Fable 5.1 Claude Code max
37.5% ±9.9
2026-09-01 431.2M $746.48
5 Grok 4.7 Grok Build xhigh
25% ±8.8
2026-09-21 437.3M $314.11
6 Fable 5 Claude Code max
25% ±8.8
2026-06-09 602.6M $1.62K
7 GPT-5.6 Sol Codex max
20.8% ±8.3
2026-07-09 1.2B $584.56
8 Gemini 3.8 Flash mini-SWE-agent high
12.5% ±6.8
2026-09-02 1.1B $117.58
9 Kimi K3 Claude Code max
12.5% ±6.8
2026-07-16 380.1M $342.60
10 GPT-5.6 Terra Codex max
8.3% ±5.6
2026-07-09 1.2B $302.41
11 DeepSeek V4.1 Flash Codex max
8.3% ±5.6
2026-09-10 1.8B $39.17
12 Grok 4.6 Grok Build xhigh
4.2% ±4.1
2026-08-12 244.9M $217.33
13 Opus 4.8 Claude Code max
4.2% ±4.1
2026-05-28 753.7M $694.26
14 GLM 5.3 Claude Code max
4.2% ±4.1
2026-08-14 1.6B $1.4K
15 GPT-5.6 Luna Codex max
0%
2026-07-09 1.6B $41.41
16 Gemini 3.7 Flash mini-SWE-agent high
0%
2026-08-13 503.6M $63.51
17 DeepSeek V4 Pro Codex max
0%
2026-08-13 1.5B $71.82
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 GPT-6 Astra Codex max
70.4% ±8.8
2026-09-03 275.5M $698.22
2 Opus 5.5 Claude Code max
63% ±9.3
2026-09-22 1.7B $1.03K
3 Fable 5.1 Claude Code max
48.1% ±9.6
2026-09-01 633.4M $976.65
4 Opus 5 Claude Code max
29.6% ±8.8
2026-07-24 1.2B $1.04K
5 Grok 4.7 Grok Build xhigh
25.9% ±8.4
2026-09-21 464.6M $343.58
6 DeepSeek V4.1 Flash Codex max
18.5% ±7.5
2026-09-10 3B $69.75
7 Grok 4.6 Grok Build xhigh
14.8% ±6.8
2026-08-12 276.7M $231.57
8 GPT-5.6 Sol Codex max
14.8% ±6.8
2026-07-09 569.2M $315.06
9 Gemini 3.8 Flash mini-SWE-agent high
11.1% ±6
2026-09-02 520.2M $70.26
10 DeepSeek V4 Pro Codex max
11.1% ±6
2026-08-13 2.2B $108.40
11 Fable 5 Claude Code max
11.1% ±6
2026-06-09 1.3B $2.39K
12 GPT-5.6 Luna Codex max
7.4% ±5
2026-07-09 1.2B $35.82
13 Gemini 3.7 Flash mini-SWE-agent high
7.4% ±5
2026-08-13 377.2M $56.41
14 GPT-5.6 Terra Codex max
7.4% ±5
2026-07-09 464.1M $141.00
15 GLM 5.3 Claude Code max
7.4% ±5
2026-08-14 992.1M $788.27
16 Opus 4.8 Claude Code max
7.4% ±5
2026-05-28 1.1B $831.49
17 Kimi K3 Claude Code max
3.7% ±3.6
2026-07-16 536.3M $354.73
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 GPT-6 Astra Codex max
80.4% ±5.6
2026-09-03 661.6M $1.45K
2 Opus 5.5 Claude Code max
76.5% ±5.9
2026-09-22 1.7B $1.16K
3 Fable 5.1 Claude Code max
41.2% ±6.9
2026-09-01 891.4M $1.6K
4 Fable 5 Claude Code max
33.3% ±6.6
2026-06-09 2.1B $4.63K
5 GPT-5.6 Sol Codex max
31.4% ±6.5
2026-07-09 3.3B $1.66K
6 Opus 5 Claude Code max
25.5% ±6.1
2026-07-24 1.7B $1.63K
7 DeepSeek V4.1 Flash Codex max
11.8% ±4.5
2026-09-10 2.9B $93.97
8 GPT-5.6 Terra Codex max
9.8% ±4.2
2026-07-09 1.6B $450.66
9 GPT-5.6 Luna Codex max
7.8% ±3.8
2026-07-09 3.7B $105.62
10 Opus 4.8 Claude Code max
7.8% ±3.8
2026-05-28 2.2B $1.9K
11 DeepSeek V4 Pro Codex max
5.9% ±3.3
2026-08-13 2.7B $146.13
12 GLM 5.3 Claude Code max
5.9% ±3.3
2026-08-14 2.1B $1.57K
13 Grok 4.6 Grok Build xhigh
5.9% ±3.3
2026-08-12 2.2B $2.07K
14 Grok 4.7 Grok Build xhigh
3.9% ±2.7
2026-09-21 1.9B $1.36K
15 Gemini 3.7 Flash mini-SWE-agent high
2% ±1.9
2026-08-13 5.7B $585.39
16 Kimi K3 Claude Code max
2% ±1.9
2026-07-16 863.2M $714.76
17 Gemini 3.8 Flash mini-SWE-agent high
0%
2026-09-02 5.2B $510.98

TAXONOMY

Domain areas

The v0.1 release spans five domain areas, unevenly weighted toward life and physical sciences. Resolution rates above are aggregated across all 70 tasks.

19

Life sciences

17

Physical sciences

8

Earth sciences

17

Mathematical sciences

9

Engineering sciences

Sample tasks

A sample of task titles from the v0.1 release. Browse the full contribution pipeline on the task dashboard →

CHEMISTRY

Constraint-Consistent Internal-Coordinate Embedding for RDKit

Implement constrained RDKit conformer generation under distance, angle, signed-torsion, and ambiguous-distance restraints.

(opens in a new tab)

ASTRONOMY

Multifrequency CMB Cross-Spectrum Inference and Blinded Model Extension
Recover CMB BB cross-spectra and tensor-to-scalar-ratio constraints, diagnose model misspecification in a blinded spectrum ensemble, and design a parametric extension that removes its cosmological bias.
(opens in a new tab)

BIOLOGY

Genomic model ranking

Rank genomic prediction models for transfer across sequence families and cell contexts, and produce calibrated probabilities for unlabelled target examples.
(opens in a new tab)

MEDICINE

Tumour-immune Spatial Interface Analysis of an Anti-PD-1 Melanoma IMC Cohort
Identify the cell populations of a melanoma imaging-mass-cytometry cohort from marker profiles, reconstruct per-ROI tumour territories from the segmentation masks, and quantify immune-cell signed-distance profiles and responder versus non-responder infiltration differences.
(opens in a new tab)

APPLIED MATHEMATICS

Koopman-Based Identification of Mean Field Game Parameters
Identify parametric and nonparametric mean field game dynamics from noisy equilibrium trajectories, construct a controlled Koopman generator, and predict held-out equilibria.
(opens in a new tab)

FORMAL MATHEMATICS

Finite free Stam inequality (GVSS Theorem 1.4)
Prove the simple-root case of the finite free Stam inequality in Lean 4.
(opens in a new tab)

GEOSCIENCES

Greenland supraglacial lake drainage classification
Classify 40 Greenland supraglacial lakes by drainage mechanism from Sentinel-2 melt-season time series, applying a written expert labeling protocol.
(opens in a new tab)

MECHANICAL ENGINEERING

ABAQUS-Informed Digital Twin for Crack Identification
Use an ABAQUS-informed digital twin and its finite-element response to localize a real crack in an experimental inspection.
(opens in a new tab)
Review pipeline

How tasks are graded

Every task in Terminal-Bench-Science goes through a review pipeline before it is eligible to be merged into a release, and every evaluation run separates the agent’s working environment from the environment that grades it.

01

Proposal

A domain expert or contributor proposes a task drawn from real research work.
02

LLM Judge

An automated judge screens the proposal for basic feasibility and completeness.
03

Reviewer

A first-pass human reviewer checks instruction clarity and task scope.
04

Senior Reviewer

A senior reviewer checks instruction and verifier alignment, task realism, and under or over specification risk.
05

Pull Request

The task is opened as a pull request against the public task repository.
06

Static Checks

Automated checks validate task structure, environment definitions, and verifier code.

07

Agent Judge

Frontier, oracle, and cheating-agent trial runs confirm the task is solvable, gradeable, and resistant to shortcuts.

08

Agent Trials

Additional agent trial runs stress test the task before it is accepted.

09

Merge

The task is merged into the public task set and becomes eligible for a future release.
of

Behind the benchmark

Terminal-Bench-Science evaluates agents on real scientific research workflows rather than software repository tasks. Version 0.1, the benchmark's first public release, covers 70 tasks across five domain areas, each sourced from work a practicing researcher actually performed.

Acknowledgments

Terminal-Bench-Science is a collaboration between Stanford University, the Laude Institute, and domain experts across scientific disciplines and research institutions worldwide, led by Steven Dillmann (Stanford University), Sanmi Koyejo (Stanford University), and Ludwig Schmidt (Stanford University, Anthropic)

Snorkel supports Terminal-Bench Science through the Open Benchmarks Grants program.

FAQs

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Opus 5.5
40.4%
3
Image
Fable 5.1
39.6%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
Grok 4.7
12.1%
3
Image
GLM 5.3
11.9%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.