Software Engineering
Open Benchmarks Grants

LibraryDesignBench

A benchmark measuring whether AI agents can design software libraries that other agents can use. One agent designs a library from an open-ended spec; three implementer agents then solve problems with it. Libraries are judged only by how simple and correct the downstream code becomes.

Built with
Image
Image
Image
Image

At a glance

15

library-design tasks

242

downstream problems

4

languages: Rust, Python, TypeScript, Haskell

11

model/harness runs evaluated

Leaderboard

Rank Model Harness Score Pass Rate Simplicity Library Cost Cost / Problem
1 Opus 5.5 mini-SWE
48.9 ±0.9
86.6% 64.5 $9.62 $0.174
2 Fable 5.1 mini-SWE
47.5 ±1.1
86.1% 62.7 $14.81 $0.198
3 GPT-6 Astra mini-SWE
45.1 ±0.7
85.7% 58.7 $3.63 $0.155
4 GPT-6 Astra Codex
44.1 ±0.8
85.5% 58.5 $4.51 $0.194
5 Kimi K3 mini-SWE
44 ±1
86% 58.6 $8.7 $0.176
6 Fable 5.1 Claude Code
42.6 ±2.5
85.8% 58.2 $14.39 $0.203
7 GLM 5.3 mini-SWE
41.9 ±1.1
84.1% 57.1 $21.06 $0.234
8 Grok 4.6 mini-SWE
39.7 ±1.1
84.5% 54.1 $2.63 $0.225
9 GPT-5.6 Sol Codex
39.5 ±1
84.1% 52.9 $2.14 $0.2
10 GPT-6 Sol mini-SWE
38.5 ±0.7
85.2% 51.6 $0.3 $0.15
11 DeepSeek V4 Pro mini-SWE
31.2 ±0.8
84.9% 42 $0.31 $0.175

Key takeaways

Opus 5.5 leads LibraryDesignBench at 48.9, ahead of Fable 5.1 at 47.5 and GPT-6 Astra at 45.1. Opus 5.5 and Fable 5.1 are the only models that beat the human-written production libraries (46.6). Ten of the eleven entries score above having no library (34.4); DeepSeek V4 Pro's libraries score below it (31.2).

Frontier performance

  • Opus 5.5
  • Fable 5.1
  • GPT-6 Astra
  • Kimi K3
  • GLM 5.3
  • Grok 4.6
  • GPT-5.6 Sol
  • GPT-6 Sol
  • DeepSeek V4 Pro
Opus 5.5 ×
Fable 5.1 ×
GPT-6 Astra ×
Kimi K3 ×
GLM 5.3 ×
Grok 4.6 ×
GPT-5.6 Sol ×
GPT-6 Sol ×
DeepSeek V4 Pro ×
Loading chart data...

Why measure library design?

Agents increasingly build on libraries other agents wrote. A good library makes the next agent's code shorter and simpler; a bad one makes it longer, or worse than writing without one. LibraryDesignBench tests the design step directly, without scoring the library's own code, size or tests.

What the results show

Agents mostly copy existing designs

On 11 of 15 tasks, designers reproduce the production library's design. In the CLI-parser task, they copy clap's builder style rather than its shorter derive macro.

A library can make things worse

DeepSeek V4 Pro's libraries score 9.2% below having no library at all. In Haskell, an agent-written library scores below no library on 70% of designer and task pairs.

The harness matters

Fable 5.1 scores 47.5 in mini-SWE-agent but 42.6 in Claude Code.

No home-team advantage

All three implementers agree on the top three designers and the last one, and none ranks its own model family higher.

Methodology

Task

Each task gives the model an open-ended spec: the capabilities the library should have and a few example uses. There are no interfaces, method signatures or tests.

Implementers

Three fixed implementer agents solve each of the task's problems using the library, without seeing the tests.

Score

Pass rate² × simplicity, averaged over every problem and implementer, with tasks weighted equally. Simplicity compares each solution with a reference solution built on the real production library. Scores run 0–100 and ± is the 95% confidence interval.

Reference arms

"No library" and "Production library" use the same implementers and problems.

Limits

No internet access. Designers get 4 hours per library. Implementers get 1 hour and $2.50 per problem. Unfinished solutions are scored as they are, which affects 1.7% of runs.

Not measured

The library's own correctness, security, runtime speed or maintainability.

Behind the benchmark

Each task asks the model under test to design a reusable library from an open-ended spec: the capabilities it should have and a few example uses. It gets no interfaces, no method signatures and no tests. There are 15 tasks in Rust, Python, TypeScript and Haskell, from a CLI argument parser like clap to a dataframe library like pandas.

The library itself is never scored: not its code, its size or its own tests. It only has to build. It is judged only by the code other agents write with it. Three fixed implementer agents solve each of the task's problems with the library, without ever seeing the tests. Each solution is scored on how simple it is compared with a reference solution built on the real production library, weighted by how many hidden tests it passes.

So a library scores well only when it makes the code built on it short and simple. The docs cover every step, including exactly what is pre-installed for each agent.

Acknowledgments

LibraryDesignBench was led by Gabe Orlanski, a Snorkel AI research fellow, with Snorkel AI's Vincent Chen and Fred Sala among the paper's co-authors. It was co-developed with researchers at the University of Wisconsin–Madison, and supported by Snorkel AI's Open Benchmarks Grants program, DARPA, the National Science Foundation, and Prime Intellect.

FAQs

Get notified when we launch a new benchmark

Share this benchmark
 Illution Back
Illution Front

For models that need to be right. Not just good enough.