agentic Coding LEADERBOARDS

Coding benchmarks for real-world software engineering

Explore benchmarks that test coding agents on the work software engineers actually do, from navigating complex codebases and debugging failures to shipping features and maintaining code quality over time.

partners
Standford logo
Image
Image
Image
Image
Image
Image
Image
agent-le-logo
Image
Image
Cua Logo

Senior SWE-bench

A benchmark for evaluating coding agents on senior-level engineering work: building features from realistic instructions, investigating bugs that require runtime investigation, and shipping code that aligns to existing codebase conventions.

Built with
Image
Image
Image
Tasteful Solve Rate
The top-performing frontier models fail to complete tasks with senior-level correctness and taste over 70% of the time.
1
Image
Claude Fable 5
34.7%
2
Image
Claude Opus 5
34.7%
3
Image
GPT-5.5
34.7%
4
Image
Claude Opus 4.8
30.5%
Open Benchmarks Grants

Terminal-Bench 3.0

The next frontier benchmark for agent work. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.

By Resolution Rate
1
Image
Opus 5 (max) • mini-SWE-agent
43.5%
2
Image
GPT-5.6 Sol (max) • Codex
32.7%
3
Image
Fable 5 (max) • Claude Code
31.7%
Open Benchmarks Grants

Agents’ Last Exam

Long-horizon professional workflows with verifiable outcomes across 55 sub-industries. 147 public tasks of a 1,500+ task corpus, sourced and validated by 300+ industry experts.

Top Submissions (Pass Rate)
1
Image
Codex · GPT 5.6-Sol · XHigh
30.6%
2
Image
Codex · GPT 5.6-Luna · XHigh
30.3%
3
Image
Codex · GPT 5.6-Sol · Max
29.6%
Open Benchmarks Grants

SlopCode Bench

Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.

Top Models by Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3-Codex
26.02%
3
Image
GPT-5.4
23.47%
Open Benchmarks Grants

Terminal-Bench 2.1

Terminal agent evaluation led by Stanford University and Laude Institute. v2.1 fixes 28 tasks from 2.0 and introduces continuous validation.

Top Submissions
1
Image
Claude Code · Claude 5 Fable
83.8%
2
Image
Codex CLI · GPT-5.5
83.1%
3
Image
Terminus 2 · Claude 5 Fable
80.4%
Agentic coding resources

Learn more

A curated collection of blogs, research papers, and reading-group discussions on agentic coding benchmarks, covering benchmark design, realistic coding environments, model performance, and the failure modes shaping what comes next.

coding-agents-eval

Why coding agents need better data, evals, and environments

Coding agents have moved from tab-complete to teammate. They autonomously inspect repositories, edit files, run commands, diagnose failures, and work through multi-step engineering tasks. That creates a harder reliability problem.
May 11, 2026
senior-swe-reading-group

Senior SWE-Bench: Evaluating Coding Agents Like Senior Engineers

At our latest Snorkel AI Reading Group, Henry Ehrenberg presented Senior SWE-Bench, an open-source, Harbor-compatible benchmark for evaluating coding agents on realistic, senior-level software engineering work. Its 100 tasks, with
July 16, 2026
Team Snorkel
Building the Benchmark Factory Banner

Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory

To kick off our inaugural Benchtalks, a series dedicated to the researchers building these measurement toolkits, Snorkel AI co-founder Vincent Sunn Chen sat down with Alex Shaw, Founding MTS at
March 31, 2026
Image

Claude Opus 5: Performance and Error Analysis on Frontier Coding Tasks

Anthropic’s Claude Opus 5 recently debuted as the second model overall on the current Senior SWE-bench leaderboard, behind Fable 5. It also achieves the highest score of any evaluated model
July 27, 2026
Closing the Evaluation Gap in Agentic AI Image

Closing the Evaluation Gap in Agentic AI

Today, AI is marked by a growing asymmetry: the excitement around agentic AI is real — backed by quantitative progress on model cards and genuine leaps forward, especially in coding.
February 11, 2026
Image

SlopCodeBench: Measuring Code Erosion as Agents Iterate

SlopCodeBench reveals how AI coding agents degrade code quality over time—measuring “slop,” technical debt, and architectural erosion across iterations.
January 20, 2026
Image

The science of rubric design

Part 3 of our rubric series explains the science of rubric design. We show why rubrics should be treated like models—structured, measured, and iterated—to maximize objective alignment and inter-rater agreement.
September 11, 2025
Image

Cua-Bench: benchmarking computer-use agents on professional software

TL;DR We built a benchmark of 25 expert-authored KiCad schematic-editing tasks and ran a frontier computer-use agent against them. The headline numbers: 1. Why build a computer-use benchmark for electrical
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.