agentic Coding LEADERBOARDS
Coding benchmarks for real-world software engineering
Explore benchmarks that test coding agents on the work software engineers actually do, from navigating complex codebases and debugging failures to shipping features and maintaining code quality over time.
Senior SWE-bench
A benchmark for evaluating coding agents on senior-level engineering work: building features from realistic instructions, investigating bugs that require runtime investigation, and shipping code that aligns to existing codebase conventions.


Terminal-Bench 3.0
The next frontier benchmark for agent work. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.
Agents’ Last Exam
Long-horizon professional workflows with verifiable outcomes across 55 sub-industries. 147 public tasks of a 1,500+ task corpus, sourced and validated by 300+ industry experts.
SlopCode Bench
Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.
Terminal-Bench 2.1
Terminal agent evaluation led by Stanford University and Laude Institute. v2.1 fixes 28 tasks from 2.0 and introduces continuous validation.
Learn more
A curated collection of blogs, research papers, and reading-group discussions on agentic coding benchmarks, covering benchmark design, realistic coding environments, model performance, and the failure modes shaping what comes next.











