Benchmarks for what frontier AI hasn't solved
Senior SWE-bench
Evaluates coding agents on senior-level software engineering tasks.


LibraryDesignBench
Measures how well AI agents design libraries that other agents can use.
Terminal-Bench-Science
Evaluates agents on scientific workflows derived from researchers’ own work.
Terminal-Bench 4.0
Evaluates terminal agents on continuously updated software-engineering tasks.
Terminal-Bench 3.0
Evaluates terminal agents on containerized tasks across seven domains.
OSWorld 2.0
Evaluates computer-use agents on 108 long-horizon workflows.
Agents’ Last Exam
Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.
Continual Learning Bench
Measures improvement across sequential, stateful tasks.
SlopCode Bench
Measures coding-agent performance across evolving software requirements.
Open Benchmarks Grants
featured Collaborations
Reasoning
ARC-AGI-3
Tests abstract reasoning through hundreds of hand-crafted interactive games, probing whether models can explore and acquire skills in novel environments.
Legal agents
Harvey’s Long Horizon Legal Agent Benchmark
Built to evaluate and improve agent capabilities for supporting legal work.
Evaluation methods
JudgmentBench
Compares rubric-based and preference-based evaluation for judging output quality.
Get notified when we launch a new benchmark
Three core dimensions where today's benchmarks fall short



