LibraryDesignBench
A benchmark measuring whether AI agents can design software libraries that other agents can use. One agent designs a library from an open-ended spec; three implementer agents then solve problems with it. Libraries are judged only by how simple and correct the downstream code becomes.
At a glance
15
library-design tasks
242
downstream problems
4
languages: Rust, Python, TypeScript, Haskell
11
model/harness runs evaluated
Leaderboard
| Rank | Model | Harness | Score | Pass Rate | Simplicity | Library Cost | Cost / Problem |
|---|---|---|---|---|---|---|---|
| 1 | Opus 5.5 | mini-SWE |
48.9
±0.9
|
86.6% | 64.5 | $9.62 | $0.174 |
| 2 | Fable 5.1 | mini-SWE |
47.5
±1.1
|
86.1% | 62.7 | $14.81 | $0.198 |
| 3 | GPT-6 Astra | mini-SWE |
45.1
±0.7
|
85.7% | 58.7 | $3.63 | $0.155 |
| 4 | GPT-6 Astra | Codex |
44.1
±0.8
|
85.5% | 58.5 | $4.51 | $0.194 |
| 5 | Kimi K3 | mini-SWE |
44
±1
|
86% | 58.6 | $8.7 | $0.176 |
| 6 | Fable 5.1 | Claude Code |
42.6
±2.5
|
85.8% | 58.2 | $14.39 | $0.203 |
| 7 | GLM 5.3 | mini-SWE |
41.9
±1.1
|
84.1% | 57.1 | $21.06 | $0.234 |
| 8 | Grok 4.6 | mini-SWE |
39.7
±1.1
|
84.5% | 54.1 | $2.63 | $0.225 |
| 9 | GPT-5.6 Sol | Codex |
39.5
±1
|
84.1% | 52.9 | $2.14 | $0.2 |
| 10 | GPT-6 Sol | mini-SWE |
38.5
±0.7
|
85.2% | 51.6 | $0.3 | $0.15 |
| 11 | DeepSeek V4 Pro | mini-SWE |
31.2
±0.8
|
84.9% | 42 | $0.31 | $0.175 |
Key takeaways
Opus 5.5 leads LibraryDesignBench at 48.9, ahead of Fable 5.1 at 47.5 and GPT-6 Astra at 45.1. Opus 5.5 and Fable 5.1 are the only models that beat the human-written production libraries (46.6). Ten of the eleven entries score above having no library (34.4); DeepSeek V4 Pro's libraries score below it (31.2).
Frontier performance
- Opus 5.5
- Fable 5.1
- GPT-6 Astra
- Kimi K3
- GLM 5.3
- Grok 4.6
- GPT-5.6 Sol
- GPT-6 Sol
- DeepSeek V4 Pro
Why measure library design?
Agents increasingly build on libraries other agents wrote. A good library makes the next agent's code shorter and simpler; a bad one makes it longer, or worse than writing without one. LibraryDesignBench tests the design step directly, without scoring the library's own code, size or tests.
What the results show
Agents mostly copy existing designs
On 11 of 15 tasks, designers reproduce the production library's design. In the CLI-parser task, they copy clap's builder style rather than its shorter derive macro.
A library can make things worse
DeepSeek V4 Pro's libraries score 9.2% below having no library at all. In Haskell, an agent-written library scores below no library on 70% of designer and task pairs.
The harness matters
Fable 5.1 scores 47.5 in mini-SWE-agent but 42.6 in Claude Code.
No home-team advantage
All three implementers agree on the top three designers and the last one, and none ranks its own model family higher.
All 15 tasks · 242 problems
Methodology
Task
Each task gives the model an open-ended spec: the capabilities the library should have and a few example uses. There are no interfaces, method signatures or tests.
Implementers
Three fixed implementer agents solve each of the task's problems using the library, without seeing the tests.
Score
Pass rate² × simplicity, averaged over every problem and implementer, with tasks weighted equally. Simplicity compares each solution with a reference solution built on the real production library. Scores run 0–100 and ± is the 95% confidence interval.
Reference arms
"No library" and "Production library" use the same implementers and problems.
Limits
No internet access. Designers get 4 hours per library. Implementers get 1 hour and $2.50 per problem. Unfinished solutions are scored as they are, which affects 1.7% of runs.
Not measured
The library's own correctness, security, runtime speed or maintainability.
Behind the benchmark
Each task asks the model under test to design a reusable library from an open-ended spec: the capabilities it should have and a few example uses. It gets no interfaces, no method signatures and no tests. There are 15 tasks in Rust, Python, TypeScript and Haskell, from a CLI argument parser like clap to a dataframe library like pandas.
The library itself is never scored: not its code, its size or its own tests. It only has to build. It is judged only by the code other agents write with it. Three fixed implementer agents solve each of the task's problems with the library, without ever seeing the tests. Each solution is scored on how simple it is compared with a reference solution built on the real production library, weighted by how many hidden tests it passes.
So a library scores well only when it makes the code built on it short and simple. The docs cover every step, including exactly what is pre-installed for each agent.
Acknowledgments
LibraryDesignBench was led by Gabe Orlanski, a Snorkel AI research fellow, with Snorkel AI's Vincent Chen and Fred Sala among the paper's co-authors. It was co-developed with researchers at the University of Wisconsin–Madison, and supported by Snorkel AI's Open Benchmarks Grants program, DARPA, the National Science Foundation, and Prime Intellect.

