Retrieved article excerpt
Open article · Retrieved 2026-10-01T02:29:17.311203+00:00
[Paper](https://arxiv.org/abs/2609.36730)[Leaderboard](https://ldbench.com)[Code](https://github.com/SprocketLab/librarydesignbench)[@Gorlanski](https://x.com/Gorlanski)
We are fast approaching a point where the primary users of software libraries are agents, not human engineers. If agents are the users, then a library should be judged by how little code agents need to write correct programs with it, not by how it reads to a human. That is why I am excited to release [LibraryDesignBench](https://ldbench.com), a two-phase benchmark that scores an agent-written library solely by how much it helps the *future agents* that use it.
## Asking agents to write libraries for other agents
LibraryDesignBench asks agents to write effective libraries, and we intentionally give them design flexibility. We never prescribe signatures their library must expose. Nor do we ever guide them to specific abstractions or patterns they should use. We intentionally give them underspecified, vague, and ambiguous library specifications because frontier agents need to envision how future agents will use their library, what they will need, and what design will work best. LibraryDesignBench gives agents this vast freedom because the evaluation needs to serve as a test bed for understanding what patterns *agents* actually prefer, not what we think they will prefer.
## Evaluating A Library Based on How Agents Use It
A library that implements something correctly does not immediately provide any value – it needs to be correct *and* make future code simpler. Thus, the *only* way to measure a library’s quality is to observe how much future agents benefit from its design decisions. LibraryDesignBench does this in two phases:
1. **Design Phase:** The agent implements a full library from an intentionally non-prescriptive specification.
2. **Evaluation Phase:** We evaluate the library by observing multiple different implementer agents attempt to solve problems using it.
The only aspect we check in the design phase is if the library is installable in their internet-restricted environment, as we pre-install all libraries for evaluation since trying to install a library is not part of the signal we care about. We then task three agents (GPT-5.6 Luna on Codex, DeepSeek v4.1 Flash on mini-SWE-agent, and GLM 5.3 Flash on mini-SWE-agent) with solving programming problems with the library in as little code as possible. Each problem ×\times× agent ×\times× library is run in its own environment without internet access, with high reasoning, and with the library pre-installed. We use a [highly prescriptive prompt](https://github.com/SprocketLab/librarydesignbench/blob/dev/configs/prompts/library_use_inst.md) to ensure agents attempt to fully exploit the library.
We now take these solutions and score each with:
score=pass rate(y)2⋅simplicity(y)\text{score} = \text{pass rate}(y)^2 \cdot \text{simplicity}(y)score=pass rate(y)2⋅simplicity(y)
Pass rate is the percentage of tests passed, and we square it to penalize incorrectness more harshly than simplicity. Simplicity is defined as:
simplicity(y)=1∣M∣∑m∈Mmin{m(y∗)m(y), 1}\text{simplicity}(y) = \frac{1}{|\mathcal{M}|} \sum\_{m \in \mathcal{M}} \min\left\{\frac{m(y^\*)}{m(y)},\, 1\right\}simplicity(y)=∣M∣1m∈M∑min{m(y)m(y∗),1}
M\mathcal{M}M is the set of static metrics we use to compare how far off the solution written with the library is from the golden reference, y∗y^\*y∗, written with the real production library. Each ratio is capped at 1, so a solution that beats the reference gets no extra credit. The four metrics are:
- **Source Lines of Code:** how much code the agent wrote.
- **Cyclomatic Complexity:** how many branches and loops the agent needed.
- **Cognitive Complexity:** how hard the code is to follow, with extra weight on nesting.
- **Halstead Volume:** size in operators and operands, which dense one-liners can’t game.
Section 2 of the [paper](https://arxiv.org/abs/2609.36730) covers the scoring and standard errors in more detail.
## Agents Copy Human Designs But Worse
We evaluate 11 designer setups, covering 9 frontier models with some run in more than one harness, all at high reasoning. Each setup attempts each of the 15 library design tasks three times, which produces 45 libraries. The three implementer agents then use each library to solve every problem in its task. Across all 242 problems, that comes to 242 problems × 3 libraries × 3 implementers = **2,178 trials per setup**.
We compare against two baselines that use the same 2,178 trials. In the first, implementers have no library and use a separate prompt. In the second, they have the human-written production library pre-installed. Neither baseline has three designed libraries to vary, so we instead run each problem-implementer pair three times.
Design Pareto
3035404550Score
↑ Beats production library
↓ Actively impedes agents
$0.5$1$2$5$10$20Cost to design one library ($, log) →
Evaluation Pareto
3035404550Score
PL
NL
$0.1$0.15$0.2Implementer cost per problem ($, log) →
Legend and notes
OpenAIAnthropicZ.aiDeepSeekMoonshot AIxAI
NL
: **No library.** The same implementers solve each problem without any library.
PL
: **Production library.** The same implementers use the task's real production library.
: **Reference score.** Dashed at the NL and PL scores in both panels. A designed library scoring below NL actively impedes the agents using it; one scoring above PL beats the production library.
: **Pareto frontier.** The best score at or below each cost. Points off it are faded.
Scores run 0–100: pass rate² × simplicity, averaged over every problem and implementer with tasks weighted equally. Hover or tap a point for exact values.
Opus 5.5 designs libraries that help downstream agents more than the human-written production libraries do. With them, implementers pass just as many tests while writing simpler code. Fable 5.1 roughly matches production. Correctness barely separates designers, since every setup passes about the same share of tests, so the ranking comes down to how much code implementers still have to write. At the other end, DeepSeek V4 Pro’s library actively hurts: implementers do worse with it than with no library at all. The harness also matters. Fable’s libraries are more useful when designed in mini-SWE-agent than in Claude Code, and Astra’s are slightly more useful in mini-SWE-agent than in Codex.
abc In all six libraries abc Astra only abc Fable only
### GPT-6 Astra
Library 1
```
Command::new("inspect")
.arg(Arg::new("verbose")
.short('v')
.action(ArgAction::Count))
.arg(Arg::new("pids")
.value_name("PID")
.num_args(1..)
.value_parser(
ValueParser::u64()));
```
Library 2
```
Command::new("inspect")
.arg(Arg::new("verbose")
.short('v')
.action(ArgAction::Count))
.arg(Arg::new("pid")
.short('p').long("pid")
.value_parser(
ValueParser::from_str::<u32>())
.required(true));
```
Library 3
```
Command::new("inspect")
.arg(Arg::new("verbose")
.short('v')
.action(ArgAction::Count))
.arg(Arg::new("pid")
.long("pid")
.action(ArgAction::Append)
.value_parser(
ValueParser::u64()));
```
### Fable 5.1
Library 1
```
Command::new("procs")
.arg(Arg::count("verbose")
.short('v'))
.arg(Arg::option("pid")
.short('p')
.uint()
.multiple());
```
Library 2
```
Command::new("proc")
.arg(Arg::count("verbose")
.short('v').global(true))
.subcommand(Command::new(
"inspect")
.arg(Arg::positional("pid")
.int().required(true));
```
Library 3
```
Command::new("procinfo")
.arg(Arg::positional("pid")
.int()
.required(true))
.arg(Arg::count("verbose")
.short('v'));
```
**Agents converge on the same designs.** README quick-starts from three
`clirs` libraries each by GPT-6 Astra and Fable 5.1, with unrelated
arguments omitted. Both copy `clap`'s builder API rather than its
shorter derive macro. Hover a key entry to spotlight it.
Beyond harness peculiarities, the biggest trend we observed is that on 11/15 tasks, agents copied a design pattern from the corresponding production library. Surprisingly, these patterns are the exact ones the implementer agents used when given the production library. Our failure analysis indicates that the implementer agents write more complex code than needed because interfaces are either too rigid to use or require verbose code. In the former case, agents reimplement functionality the libraries provide, adding bugs in the process.
Failed tests
— why at least one test failed
GPT-6 Astra
15%
10%
18%
9%
48%
GPT-6 Sol
24%
13%
18%
5%
41%
Fable 5.1
16%
17%
19%
6%
42%
Opus 5.5
17%
13%
15%
6%
50%
GLM 5.3
14%
19%
14%
8%
44%
Grok 4.6
12%
25%
19%
12%
32%
Excess code
— why the solution was longer than the reference
GPT-6 Astra
21%
35%
30%
14%
GPT-6 Sol
14%
44%
24%
16%
Fable 5.1
16%
31%
30%
21%
Opus 5.5
13%
35%
27%
24%
GLM 5.3
12%
7%
30%
30%
21%
Grok 4.6
11%
7%
36%
33%
14%
Limited by the library:
Coverage
Correctness
Rigidity
Verbosity
Library not fully exploited:
Underused library
**Most excess code traces back to the library.** Primary failure category
per designer. **Failed tests** classifies why at least one test failed;
**excess code** classifies why the solution was longer than the reference.
Categories and method
Limited by the library: no path through the library as shipped does
better.
Coverage
: Nothing close to the needed capability exists; the fix is a new operation.
Correctness
: The capability exists but has a bug: wrong output on a legitimate input, a misleading error, or a path too slow for the problem.
Rigidity
: It nearly fits but cannot be adapted: a hard-coded policy, a data model that cannot hold the value, a monolithic operation, or a scope that excludes the input.
Verbosity
: It fits, but using it takes excess code: a verbose interface, unpackaged wiring, shape conversion, or repeated declarations.
Library not fully exploited: a path through the library as shipped removes
the code or fixes the failure.
Underused library
: A simpler path through the library exists, but the implementer never found it, did not recognize it, or missed a fact needed to use it.
**How.** For each library from six mini-SWE-agent designers, we sampled
three solutions that failed at least one test and wrote more code than the
reference. A GPT-5.6 Luna auditor read the library, the implementer's trajectory,
the tests, and the reference, then classified each symptom by the smallest
library change that would have prevented it. The bars show the primary cause:
810 classifications per symptom, pooled across designers.
**Takeaway.** 82% of excess-code causes are limits of the library itself.
Rigidity and Verbosity alone account for 64%, against 14% for missing capabilities.
For failed tests, an underused library leads at 43%, and over half of those
are an incomplete contract: the implementer found the capability but not a
precondition or default needed to use it.
Each solution was audited once by one model, without running the proposed
path. Shares describe partially passing solutions, not every solution.
## Can Agents Use Libraries Without Instruction?
Our main evaluation prompt is rather heavy-handed because we need agents to fully exploit the library to measure its upper bound. But we also want to understand how *models* will behave given [minimal additional instructions](https://github.com/SprocketLab/librarydesignbench/blob/dev/configs/prompts/minimal.md). Thus, we also release **LibraryUseBench**, our evaluation phase in which agents must use the production library to write as little code as possible. All agents use mini-SWE-agent and use High thinking. Opus 5.5, unsurprisingly, is the best model evaluated, but Sonnet 5.5 is close behind at just over half the cost per problem. The gap between them is sma