Benchmarks for what frontier AI hasn't solved
Senior SWE-bench
A benchmark for evaluating coding agents on senior-level engineering work: building features from realistic instructions, investigating bugs that require runtime investigation, and shipping code that aligns to existing codebase conventions.


Terminal-Bench 4.0
A benchmark to measure and evolve with the frontier of agent work: real terminal environments, real software engineering tasks, and a rolling task set that is revised as agents catch up to it.
Terminal-Bench-Science
A benchmark for evaluating AI agents on workflows from researchers’ own work. Scientists, not model developers or vendors, set the bar for scientific capability in AI.
Terminal-Bench 3.0
The next frontier benchmark for agent work. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.
OSWorld 2.0
A benchmark for evaluating computer-use agents on long-horizon, real-world workflows: 108 authentic tasks across 31 self-hosted web environments and professional desktop applications.
Agents’ Last Exam
Long-horizon professional workflows with verifiable outcomes across 55 sub-industries. 147 public tasks of a 1,500+ task corpus, sourced and validated by 300+ industry experts.
Continual Learning Bench
Evaluates whether AI systems improve from prior experience across sequential, stateful tasks, measuring real in-context learning, not just raw capability.
SlopCode Bench
Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.
Open Benchmarks Grants
featured Collaborations
Reasoning
ARC-AGI-3
Tests abstract reasoning through hundreds of hand-crafted interactive games, probing whether models can explore and acquire skills in novel environments.
Legal agents
Harvey’s Long Horizon Legal Agent Benchmark
Built to evaluate and improve agent capabilities for supporting legal work.
Evaluation methods
JudgmentBench
Compares rubric-based and preference-based evaluation for judging output quality.
Get notified when we launch a new benchmark
Three core dimensions where today's benchmarks fall short



