2026-10-11 16:37 UTC

Gabe Orlanski's released LibraryDesignBench claims frontier agents — Opus 5.5 above all — can design agent-facing libraries that beat human-written production libraries at pass-rate² × simplicity across downstream implementer agents, and third-party adoption of the benchmark and leaderboard (or fade and refutation of the claim) settles whether agents-as-library-users becomes a measured engineering capability.

state: seedheat: lowuncertainty: mediumconvergesscott: highagent-evaluation agent-harnesses coding-agentsGabe Orlanski

What is this?

The supplied web results contain no independent coverage of Gabe Orlanski or LibraryDesignBench — the case rests entirely on its own first-party evidence (a release with paper, leaderboard, and code, plus a paper titled 'Can Agents Design Libraries for Agents?'), and the snippets cannot corroborate the Opus 5.5 leader claim (they reference Opus 4.5–4.8 but no Opus 5.5). What the results do establish is the surrounding mid-2026 landscape: an extremely crowded agent-benchmark field (SWE-bench Verified/Pro, ALE, ITBench-AA, CHI-Bench, JetBrains' Kotlin benchmark) with a hostile credibility climate — Berkeley showed eight popular benchmarks can be gamed to near-perfect scores, SWE-bench Verified scores run 15–30 points ahead of SWE-bench Pro, and ALE finds frontier agents passing only ~26% of real professional work. A first-party benchmark claiming agents beat human-written production libraries thus lands in an environment where the default posture toward new leaderboards is skepticism, and third-party adoption — not the release itself — is the real test of the claim.

Why it matters to Scott

LibraryDesignBench independently operationalizes the Reflexive Agent Design thesis — agent-facing interfaces as a distinct design discipline, judged by replaying designs against downstream agent usage — and makes simplicity a leaderboarded scoring axis that his attention-budget/clutter-tax pages argued for without ever having a public instrument. If third parties adopt the benchmark he gets dated receipts ('I argued this; now it's measured'); if it fades or gets gamed it lands squarely in his specification-gaming / challenger-never-arbiter territory — and either resolution bears directly on the agent-facing surfaces (MCP connectors, harness tooling) he builds daily. Caveat kept honest: this is first-party convergence at thesis level, with no independent coverage of Orlanski or the Opus 5.5 leader claim, so the radar should watch adoption, not the release.
ip:framework.reflexive-agent-designip:source.reflexive-agent-design-ebookip:framework.code-first-architectureip:source.why-code-execution-beats-mcpip:framework.agent-native-computingip:concept.attention-budgetradar:harnessopt-agent-harness-optimization-benchmarkradar:devtool-agent-experience-kitradar:gauge-ax-check-agent-onboardingradar:tdqs-mcp-tool-quality-specradar:post-merge-agentic-code-benchmark
queries asked of Scott's wikis
  • designing tools and APIs for agent consumption — MCP tool description ergonomics
  • benchmark gaming, leaderboard skepticism, agent eval validity
  • agent-authored code versus production human-written code quality
  • coding agent harness design — tools written for the model not the human
  • simplicity or token-efficiency as an explicit objective in agent interfaces
  • evaluation setups where a downstream agent is the consumer or judge

Measured heat

now 0 pts/hpeak 1 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 257h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-01 02:32 (minted)⭐ origin echo-reconstructedReleases LibraryDesignBench/LibraryUseBench (paper, leaderboard, code): agents design libraries from non-prescriptive specs, scored by pass
Gabe Orlanski on blog (echo) · attributed from hn.story.49915567 · published time unknown
—
09-30 22:57first on hacker news · published · lag ?Can Agents Design Libraries for Agents?
matt_d
—
09-30 22:57amplified on hacker news 👑hn.story.49915567
matt_d
peak 1 · 0 comments · 106% of case engagement
10-01 00:21our radar first saw it · lag ?discovery anchor: hn.story.49915567—
pace: p8 vs 1188 stories at the 168h mark (now 257h old) — behind addom-local-coding-harness (0.5x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnCan Agents Design Libraries for Agents?
Retrieved article excerpt

Open article · Retrieved 2026-10-01T02:29:17.311203+00:00

[Paper](https://arxiv.org/abs/2609.36730)[Leaderboard](https://ldbench.com)[Code](https://github.com/SprocketLab/librarydesignbench)[@Gorlanski](https://x.com/Gorlanski)

We are fast approaching a point where the primary users of software libraries are agents, not human engineers. If agents are the users, then a library should be judged by how little code agents need to write correct programs with it, not by how it reads to a human. That is why I am excited to release [LibraryDesignBench](https://ldbench.com), a two-phase benchmark that scores an agent-written library solely by how much it helps the *future agents* that use it.

## Asking agents to write libraries for other agents

LibraryDesignBench asks agents to write effective libraries, and we intentionally give them design flexibility. We never prescribe signatures their library must expose. Nor do we ever guide them to specific abstractions or patterns they should use. We intentionally give them underspecified, vague, and ambiguous library specifications because frontier agents need to envision how future agents will use their library, what they will need, and what design will work best. LibraryDesignBench gives agents this vast freedom because the evaluation needs to serve as a test bed for understanding what patterns *agents* actually prefer, not what we think they will prefer.

## Evaluating A Library Based on How Agents Use It

A library that implements something correctly does not immediately provide any value – it needs to be correct *and* make future code simpler. Thus, the *only* way to measure a library’s quality is to observe how much future agents benefit from its design decisions. LibraryDesignBench does this in two phases:

1. **Design Phase:** The agent implements a full library from an intentionally non-prescriptive specification.
2. **Evaluation Phase:** We evaluate the library by observing multiple different implementer agents attempt to solve problems using it.

The only aspect we check in the design phase is if the library is installable in their internet-restricted environment, as we pre-install all libraries for evaluation since trying to install a library is not part of the signal we care about. We then task three agents (GPT-5.6 Luna on Codex, DeepSeek v4.1 Flash on mini-SWE-agent, and GLM 5.3 Flash on mini-SWE-agent) with solving programming problems with the library in as little code as possible. Each problem ×\times× agent ×\times× library is run in its own environment without internet access, with high reasoning, and with the library pre-installed. We use a [highly prescriptive prompt](https://github.com/SprocketLab/librarydesignbench/blob/dev/configs/prompts/library_use_inst.md) to ensure agents attempt to fully exploit the library.

We now take these solutions and score each with:

score=pass rate(y)2⋅simplicity(y)\text{score} = \text{pass rate}(y)^2 \cdot \text{simplicity}(y)score=pass rate(y)2⋅simplicity(y)

Pass rate is the percentage of tests passed, and we square it to penalize incorrectness more harshly than simplicity. Simplicity is defined as:

simplicity(y)=1∣M∣∑m∈Mmin⁡{m(y∗)m(y), 1}\text{simplicity}(y) = \frac{1}{|\mathcal{M}|} \sum\_{m \in \mathcal{M}} \min\left\{\frac{m(y^\*)}{m(y)},\, 1\right\}simplicity(y)=∣M∣1​m∈M∑​min{m(y)m(y∗)​,1}

M\mathcal{M}M is the set of static metrics we use to compare how far off the solution written with the library is from the golden reference, y∗y^\*y∗, written with the real production library. Each ratio is capped at 1, so a solution that beats the reference gets no extra credit. The four metrics are:

- **Source Lines of Code:** how much code the agent wrote.
- **Cyclomatic Complexity:** how many branches and loops the agent needed.
- **Cognitive Complexity:** how hard the code is to follow, with extra weight on nesting.
- **Halstead Volume:** size in operators and operands, which dense one-liners can’t game.

Section 2 of the [paper](https://arxiv.org/abs/2609.36730) covers the scoring and standard errors in more detail.

## Agents Copy Human Designs But Worse

We evaluate 11 designer setups, covering 9 frontier models with some run in more than one harness, all at high reasoning. Each setup attempts each of the 15 library design tasks three times, which produces 45 libraries. The three implementer agents then use each library to solve every problem in its task. Across all 242 problems, that comes to 242 problems × 3 libraries × 3 implementers = **2,178 trials per setup**.

We compare against two baselines that use the same 2,178 trials. In the first, implementers have no library and use a separate prompt. In the second, they have the human-written production library pre-installed. Neither baseline has three designed libraries to vary, so we instead run each problem-implementer pair three times.

Design Pareto

3035404550Score

↑ Beats production library

↓ Actively impedes agents

$0.5$1$2$5$10$20Cost to design one library ($, log) →

Evaluation Pareto

3035404550Score

PL

NL

$0.1$0.15$0.2Implementer cost per problem ($, log) →

Legend and notes

OpenAIAnthropicZ.aiDeepSeekMoonshot AIxAI

NL
:   **No library.** The same implementers solve each problem without any library.

PL
:   **Production library.** The same implementers use the task's real production library.
:   **Reference score.** Dashed at the NL and PL scores in both panels. A designed library scoring below NL actively impedes the agents using it; one scoring above PL beats the production library.
:   **Pareto frontier.** The best score at or below each cost. Points off it are faded.

Scores run 0–100: pass rate² × simplicity, averaged over every problem and implementer with tasks weighted equally. Hover or tap a point for exact values.

Opus 5.5 designs libraries that help downstream agents more than the human-written production libraries do. With them, implementers pass just as many tests while writing simpler code. Fable 5.1 roughly matches production. Correctness barely separates designers, since every setup passes about the same share of tests, so the ranking comes down to how much code implementers still have to write. At the other end, DeepSeek V4 Pro’s library actively hurts: implementers do worse with it than with no library at all. The harness also matters. Fable’s libraries are more useful when designed in mini-SWE-agent than in Claude Code, and Astra’s are slightly more useful in mini-SWE-agent than in Codex.

abc In all six libraries  abc Astra only  abc Fable only

### GPT-6 Astra

Library 1

```
Command::new("inspect")
  .arg(Arg::new("verbose")
    .short('v')
    .action(ArgAction::Count))
  .arg(Arg::new("pids")
    .value_name("PID")
    .num_args(1..)
    .value_parser(
      ValueParser::u64()));
```

Library 2

```
Command::new("inspect")
  .arg(Arg::new("verbose")
    .short('v')
    .action(ArgAction::Count))
  .arg(Arg::new("pid")
    .short('p').long("pid")
    .value_parser(
      ValueParser::from_str::<u32>())
    .required(true));
```

Library 3

```
Command::new("inspect")
  .arg(Arg::new("verbose")
    .short('v')
    .action(ArgAction::Count))
  .arg(Arg::new("pid")
    .long("pid")
    .action(ArgAction::Append)
    .value_parser(
      ValueParser::u64()));
```

### Fable 5.1

Library 1

```
Command::new("procs")
  .arg(Arg::count("verbose")
    .short('v'))
  .arg(Arg::option("pid")
    .short('p')
    .uint()
    .multiple());
```

Library 2

```
Command::new("proc")
  .arg(Arg::count("verbose")
    .short('v').global(true))
  .subcommand(Command::new(
      "inspect")
    .arg(Arg::positional("pid")
      .int().required(true));
```

Library 3

```
Command::new("procinfo")
  .arg(Arg::positional("pid")
    .int()
    .required(true))
  .arg(Arg::count("verbose")
    .short('v'));
```

**Agents converge on the same designs.** README quick-starts from three
`clirs` libraries each by GPT-6 Astra and Fable 5.1, with unrelated
arguments omitted. Both copy `clap`'s builder API rather than its
shorter derive macro. Hover a key entry to spotlight it.

 

Beyond harness peculiarities, the biggest trend we observed is that on 11/15 tasks, agents copied a design pattern from the corresponding production library. Surprisingly, these patterns are the exact ones the implementer agents used when given the production library. Our failure analysis indicates that the implementer agents write more complex code than needed because interfaces are either too rigid to use or require verbose code. In the former case, agents reimplement functionality the libraries provide, adding bugs in the process.

Failed tests

— why at least one test failed

GPT-6 Astra

15%

10%

18%

9%

48%

GPT-6 Sol

24%

13%

18%

5%

41%

Fable 5.1

16%

17%

19%

6%

42%

Opus 5.5

17%

13%

15%

6%

50%

GLM 5.3

14%

19%

14%

8%

44%

Grok 4.6

12%

25%

19%

12%

32%

Excess code

— why the solution was longer than the reference

GPT-6 Astra

21%

35%

30%

14%

GPT-6 Sol

14%

44%

24%

16%

Fable 5.1

16%

31%

30%

21%

Opus 5.5

13%

35%

27%

24%

GLM 5.3

12%

7%

30%

30%

21%

Grok 4.6

11%

7%

36%

33%

14%

Limited by the library:

Coverage

Correctness

Rigidity

Verbosity

Library not fully exploited:

Underused library

**Most excess code traces back to the library.** Primary failure category
per designer. **Failed tests** classifies why at least one test failed;
**excess code** classifies why the solution was longer than the reference.

  Categories and method

Limited by the library: no path through the library as shipped does
better.

Coverage
:   Nothing close to the needed capability exists; the fix is a new operation.

Correctness
:   The capability exists but has a bug: wrong output on a legitimate input, a misleading error, or a path too slow for the problem.

Rigidity
:   It nearly fits but cannot be adapted: a hard-coded policy, a data model that cannot hold the value, a monolithic operation, or a scope that excludes the input.

Verbosity
:   It fits, but using it takes excess code: a verbose interface, unpackaged wiring, shape conversion, or repeated declarations.

Library not fully exploited: a path through the library as shipped removes
the code or fixes the failure.

Underused library
:   A simpler path through the library exists, but the implementer never found it, did not recognize it, or missed a fact needed to use it.

**How.** For each library from six mini-SWE-agent designers, we sampled
three solutions that failed at least one test and wrote more code than the
reference. A GPT-5.6 Luna auditor read the library, the implementer's trajectory,
the tests, and the reference, then classified each symptom by the smallest
library change that would have prevented it. The bars show the primary cause:
810 classifications per symptom, pooled across designers.

**Takeaway.** 82% of excess-code causes are limits of the library itself.
Rigidity and Verbosity alone account for 64%, against 14% for missing capabilities.
For failed tests, an underused library leads at 43%, and over half of those
are an incomplete contract: the implementer found the capability but not a
precondition or default needed to use it.

Each solution was audited once by one model, without running the proposed
path. Shares describe partially passing solutions, not every solution.

 

## Can Agents Use Libraries Without Instruction?

Our main evaluation prompt is rather heavy-handed because we need agents to fully exploit the library to measure its upper bound. But we also want to understand how *models* will behave given [minimal additional instructions](https://github.com/SprocketLab/librarydesignbench/blob/dev/configs/prompts/minimal.md). Thus, we also release **LibraryUseBench**, our evaluation phase in which agents must use the production library to write as little code as possible. All agents use mini-SWE-agent and use High thinking. Opus 5.5, unsurprisingly, is the best model evaluated, but Sonnet 5.5 is close behind at just over half the cost per problem. The gap between them is sma
matt_d10
🟧 echo.blog ⭐Releases LibraryDesignBench/LibraryUseBench (paper, leaderboard, code): agents design libraries from non-prescriptive specs, scored by pass Gabe Orlanski——

Interpretation history

Decision trace