Open Benchmarks Grants

Terminal-Bench-Science

Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work. Scientists, not model developers or vendors, set the bar for scientific capability in AI.

Built with
Image
Image
Image
Image

Overview

Terminal-Bench-Science evaluates agents on real scientific research workflows rather than software repository tasks. Version 0.1, the benchmark’s first public release, covers 70 tasks across five domain areas, each sourced from work a practicing researcher actually performed.

Headline finding: the strongest model evaluated, Claude Opus 5, resolves 30.0% of tasks in this release, well below the 50 to 80% range frontier models reach on general coding benchmarks. In this evaluation, authentic scientific workflows remain considerably harder for agents than general software engineering.

Snorkel supports Terminal-Bench Science through the Open Benchmarks Grants program.

At a glance

70

tasks in the v0.1 release

5

domain areas covered

9

models evaluated

30.0%

top resolution rate (Claude Opus 5)
 

Leaderboard

Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Opus 5 Claude Code max
30% ±3.2
2026-07-24 7.3B $6.99K
2 GPT-5.6 Sol Codex max
22.4% ±2.9
2026-07-09 8.4B $4.22K
3 Claude Fable 5 Claude Code max
21.4% ±2.8
2026-06-09 6.4B $14.18K
4 Claude Opus 4.8 Claude Code max
10.5% ±2.1
2026-05-28 6.8B $5.84K
5 GPT-5.6 Terra Codex max
8.6% ±1.9
2026-07-09 7.6B $1.98K
6 GLM 5.3 Claude Code max
8.1% ±1.9
2026-08-14 8.5B $2.73K
7 Kimi K3 Claude Code max
7.1% ±1.8
2026-07-16 3.2B $1.52K
8 Grok 4.6 Grok Build xhigh
7.1% ±1.8
2026-08-12 3.7B $3.34K
9 GPT-5.6 Luna Codex max
3.3% ±1.2
2026-07-09 14.2B $383.20
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Opus 5 Claude Code max
29.8%
2026-07-24 7.3B $6.99K
2 GPT-5.6 Sol Codex max
17.5%
2026-07-09 8.4B $4.22K
3 Claude Fable 5 Claude Code max
17.5%
2026-06-09 6.4B $14.18K
4 GLM 5.3 Claude Code max
10.5%
2026-08-14 8.5B $2.73K
5 Claude Opus 4.8 Claude Code max
7%
2026-05-28 6.8B $5.84K
6 GPT-5.6 Terra Codex max
7%
2026-07-09 7.6B $1.98K
7 Kimi K3 Claude Code max
7%
2026-07-16 3.2B $1.52K
8 Grok 4.6 Grok Build xhigh
7%
2026-08-12 3.7B $3.34K
9 GPT-5.6 Luna Codex max
1.8%
2026-07-09 14.2B $383.20
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Opus 5 Claude Code max
45.8%
2026-07-24 7.3B $6.99K
2 Claude Fable 5 Claude Code max
25%
2026-06-09 6.4B $14.18K
3 GPT-5.6 Sol Codex max
20.8%
2026-07-09 8.4B $4.22K
4 Kimi K3 Claude Code max
12.5%
2026-07-16 3.2B $1.52K
5 GPT-5.6 Terra Codex max
8.3%
2026-07-09 7.6B $1.98K
6 Claude Opus 4.8 Claude Code max
4.2%
2026-05-28 6.8B $5.84K
7 GLM 5.3 Claude Code max
4.2%
2026-08-14 8.5B $2.73K
8 Grok 4.6 Grok Build xhigh
4.2%
2026-08-12 3.7B $3.34K
9 GPT-5.6 Luna Codex max
0%
2026-07-09 14.2B $383.20
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Opus 5 Claude Code max
29.6%
2026-07-24 7.3B $6.99K
2 GPT-5.6 Sol Codex max
14.8%
2026-07-09 8.4B $4.22K
3 Grok 4.6 Grok Build xhigh
14.8%
2026-08-12 3.7B $3.34K
4 Claude Fable 5 Claude Code max
11.1%
2026-06-09 6.4B $14.18K
5 Claude Opus 4.8 Claude Code max
7.4%
2026-05-28 6.8B $5.84K
6 GPT-5.6 Terra Codex max
7.4%
2026-07-09 7.6B $1.98K
7 GLM 5.3 Claude Code max
7.4%
2026-08-14 8.5B $2.73K
8 GPT-5.6 Luna Codex max
7.4%
2026-07-09 14.2B $383.20
9 Kimi K3 Claude Code max
3.7%
2026-07-16 3.2B $1.52K
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Fable 5 Claude Code max
33.3%
2026-06-09 6.4B $14.18K
2 GPT-5.6 Sol Codex max
31.4%
2026-07-09 8.4B $4.22K
3 Claude Opus 5 Claude Code max
25.5%
2026-07-24 7.3B $6.99K
4 GPT-5.6 Terra Codex max
9.8%
2026-07-09 7.6B $1.98K
5 Claude Opus 4.8 Claude Code max
7.8%
2026-05-28 6.8B $5.84K
6 GPT-5.6 Luna Codex max
7.8%
2026-07-09 14.2B $383.20
7 GLM 5.3 Claude Code max
5.9%
2026-08-14 8.5B $2.73K
8 Grok 4.6 Grok Build xhigh
5.9%
2026-08-12 3.7B $3.34K
9 Kimi K3 Claude Code max
2%
2026-07-16 3.2B $1.52K
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Opus 5 Claude Code max
27.5%
2026-07-24 7.3B $6.99K
2 GPT-5.6 Sol Codex max
23.5%
2026-07-09 8.4B $4.22K
3 Claude Opus 4.8 Claude Code max
21.6%
2026-05-28 6.8B $5.84K
4 Claude Fable 5 Claude Code max
17.6%
2026-06-09 6.4B $14.18K
5 Kimi K3 Claude Code max
11.8%
2026-07-16 3.2B $1.52K
6 GPT-5.6 Terra Codex max
9.8%
2026-07-09 7.6B $1.98K
7 GLM 5.3 Claude Code max
9.8%
2026-08-14 8.5B $2.73K
8 Grok 4.6 Grok Build xhigh
5.9%
2026-08-12 3.7B $3.34K
9 GPT-5.6 Luna Codex max
0%
2026-07-09 14.2B $383.20

Terminal-Bench-Science 0.1 Pareto Frontier

Loading chart data...

TAXONOMY

Domain areas

The v0.1 release spans five domain areas, unevenly weighted toward life and physical sciences. Resolution rates above are aggregated across all 70 tasks.

19

Life sciences

17

Physical sciences

8

Earth sciences

17

Mathematical sciences

9

Engineering sciences

Sample tasks

A sample of task titles from the v0.1 release. Browse the full contribution pipeline on the task dashboard →

CHEMISTRY

Constraint-Consistent Internal-Coordinate Embedding for RDKit

Implement constrained RDKit conformer generation under distance, angle, signed-torsion, and ambiguous-distance restraints.

ASTRONOMY

Multifrequency CMB Cross-Spectrum Inference and Blinded Model Extension
Recover CMB BB cross-spectra and tensor-to-scalar-ratio constraints, diagnose model misspecification in a blinded spectrum ensemble, and design a parametric extension that removes its cosmological bias.

BIOLOGY

Genomic model ranking

Rank genomic prediction models for transfer across sequence families and cell contexts, and produce calibrated probabilities for unlabelled target examples.

MEDICINE

Tumour-immune Spatial Interface Analysis of an Anti-PD-1 Melanoma IMC Cohort
Identify the cell populations of a melanoma imaging-mass-cytometry cohort from marker profiles, reconstruct per-ROI tumour territories from the segmentation masks, and quantify immune-cell signed-distance profiles and responder versus non-responder infiltration differences.

APPLIED MATHEMATICS

Koopman-Based Identification of Mean Field Game Parameters
Identify parametric and nonparametric mean field game dynamics from noisy equilibrium trajectories, construct a controlled Koopman generator, and predict held-out equilibria.

FORMAL MATHEMATICS

Finite free Stam inequality (GVSS Theorem 1.4)
Prove the simple-root case of the finite free Stam inequality in Lean 4.

GEOSCIENCES

Greenland supraglacial lake drainage classification
Classify 40 Greenland supraglacial lakes by drainage mechanism from Sentinel-2 melt-season time series, applying a written expert labeling protocol.

MECHANICAL ENGINEERING

ABAQUS-Informed Digital Twin for Crack Identification
Use an ABAQUS-informed digital twin and its finite-element response to localize a real crack in an experimental inspection.
Review pipeline

How tasks are graded

Every task in Terminal-Bench-Science goes through a review pipeline before it is eligible to be merged into a release, and every evaluation run separates the agent’s working environment from the environment that grades it.

01

Proposal

A domain expert or contributor proposes a task drawn from real research work.
02

LLM Judge

An automated judge screens the proposal for basic feasibility and completeness.
03

Reviewer

A first-pass human reviewer checks instruction clarity and task scope.
04

Senior Reviewer

A senior reviewer checks instruction and verifier alignment, task realism, and under or over specification risk.
05

Pull Request

The task is opened as a pull request against the public task repository.
06

Static Checks

Automated checks validate task structure, environment definitions, and verifier code.

07

Agent Judge

Frontier, oracle, and cheating-agent trial runs confirm the task is solvable, gradeable, and resistant to shortcuts.

08

Agent Trials

Additional agent trial runs stress test the task before it is accepted.

09

Merge

The task is merged into the public task set and becomes eligible for a future release.
of

Acknowledgments

Terminal-Bench-Science is a collaboration between Stanford University, the Laude Institute, and domain experts across scientific disciplines and research institutions worldwide, led by Steven Dillmann (Stanford University), Sanmi Koyejo (Stanford University), and Ludwig Schmidt (Stanford University, Anthropic)

Snorkel supports Terminal-Bench Science through the Open Benchmarks Grants program.

FAQs

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

Open Benchmarks Grants

Terminal-Bench 3.0

The next frontier benchmark for agent work. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.

By Resolution Rate
1
Image
Opus 5 (max) • mini-SWE-agent
42.7%
2
Image
GPT-5.6 Sol (max) • Codex
34.6%
3
Image
Fable 5 (max) • Claude Code
34.1%

Senior SWE-bench

A benchmark for evaluating coding agents on senior-level engineering work: building features from realistic instructions, investigating bugs that require runtime investigation, and shipping code that aligns to existing codebase conventions.

Tasteful Solve Rate
1
Image
Claude Fable 5
34.7%
2
Image
Claude Opus 5
34.7%
3
Image
GPT-5.6 Sol
34.7%
Open Benchmarks Grants

OSWorld 2.0

A benchmark for evaluating computer-use agents on long-horizon, real-world workflows: 108 authentic tasks across 31 self-hosted web environments and professional desktop applications.

By Binary Accuracy (300 steps)
1
Image
GPT-5.5 · xhigh
13%
2
Image
Claude Opus 4.7 · max
13%
3
Image
Claude Sonnet 4.6 · medium
8.3%
Open Benchmarks Grants

Agents’ Last Exam

Long-horizon professional workflows with verifiable outcomes across 55 sub-industries. 147 public tasks of a 1,500+ task corpus, sourced and validated by 300+ industry experts.

Top Submissions (Pass Rate)
1
Image
Codex · GPT 5.6-Sol · XHigh
30.6%
2
Image
Codex · GPT 5.6-Luna · XHigh
30.3%
3
Image
Kimi Code · Kimi K3 · Max
28.3%
Open Benchmarks Grants

SlopCode Bench

Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.

Top Models by Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3-Codex
26.02%
3
Image
GPT-5.4
23.47%
Open Benchmarks Grants

Continual Learning Bench

Evaluates whether AI systems improve from prior experience across sequential, stateful tasks, measuring real in-context learning, not just raw capability.

Top Systems (Agg. Reward)
1
Image
ICL · Claude Sonnet 4.6
+0.196
2
Image
ICL · GPT-5.4
+0.189
3
Image
Claude Code · Sonnet 4.6
+0.185
Open Benchmarks Grants

Terminal-Bench 2.1

Terminal agent evaluation led by Stanford University and Laude Institute. v2.1 fixes 28 tasks from 2.0 and introduces continuous validation.

Top Submissions
1
Image
Claude Code · Claude 5 Fable
83.8%
2
Image
Codex CLI · GPT-5.5
83.1%
3
Image
Terminus 2 · Claude 5 Fable
80.4%
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.