2026-10-11 16:37 UTC

Apollo GraphQL's published benchmark claims GraphQL-backed MCP tools complete its tested Haiku-and-Goose tasks at lower token usage and inference cost than REST-backed alternatives, potentially making server-side joins and field selection material agent-interface optimizations.

state: seedheat: mediumuncertainty: mediumconvergesscott: mediummcp agent-harnesses inference-economicsApollo GraphQL

What is this?

The supplied case describes an Apollo GraphQL benchmark comparing GraphQL-backed and REST-backed MCP tools for agent tasks. Its evidence titles report 264 runs using Claude Haiku 4.5 through Goose, with median best-cell REST-to-GraphQL ratios of 3.19× for inference cost and 4.49× for token usage. No web results were returned, so these remain case-reported claims: the supplied material does not establish the benchmark methodology, comparable task success, or whether server-side joins and field selection caused the reported advantage.

Why it matters to Scott

Apollo’s reported comparison offers a concrete test of Scott’s Code-First Architecture claim that interface design and intermediate-data traffic impose a context tax, potentially extending its optimization options to GraphQL-backed MCP tools for projects such as MCP IP Wiki. This warrants an equal-success, trace-backed comparison rather than a migration: the supplied evidence establishes neither comparable task success nor the proposed joins/field-selection mechanism, and the radar’s related composition and harness-cost cases do not already track this Apollo benchmark.
ip:framework.code-first-architectureip:concept.signal-extractiondev:project.mcp-ip-wikidev:concept.trace-backed-agent-comparisonradar:countinghouse-in-process-mcp-compositionradar:frontierharness-17x-cost-variation
queries asked of Scott's wikis
  • MCP tool interface design REST GraphQL
  • agent harness benchmarks task success token cost
  • server-side joins versus agent tool-call orchestration
  • field selection response shaping context efficiency
  • inference economics tool schema overhead

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 578h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-17 14:22 (minted)⭐ origin echo-reconstructedReports 264 runs using Claude Haiku 4.5 through Goose, with median best-cell REST-to-GraphQL ratios of 3.19× on cost and 4.49× on tokens in
Apollo GraphQL on github (echo) · attributed from hn.story.49741162 · published time unknown
—
09-17 14:18first on hacker news · published · lag ?GraphQL-backed MCP tools are more token-efficient
jdauriemma
—
09-17 14:18amplified on hacker news 👑hn.story.49741162
jdauriemma
peak 2 · 0 comments · 98% of case engagement
09-17 14:20our radar first saw it · lag ?discovery anchor: hn.story.49741162—
pace: p9 vs 1032 stories at the 336h mark (now 578h old) — behind addom-local-coding-harness (0.5x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnGraphQL-backed MCP tools are more token-efficient
Retrieved article excerpt

Open article · Retrieved 2026-09-17T14:22:18.686993+00:00

# GraphQL-backed MCP tools are more token-efficient

This benchmark suite evaluates a variety of agentic tasks run against multiple different setups, some backed by a GraphQL API and some backed by a REST API.  GraphQL-over-MCP tool calls achieve task success using fewer tokens than REST-over-MCP and with lower overall inference costs.  I am comparing apples-to-apples: even REST API variants implemented with OpenAPI schema and field selection capabilities do not match GraphQL’s token efficiency.  The results favor GraphQL for both trivial tasks that query a real production system (GitHub) and tasks that require fetching from multiple services/entities (a mocked travel-booking API, the GraphQL variant of which uses Federation for orchestration). Three mechanisms drive that gap: field selection is the default grammar of a GraphQL query rather than an opt-in bracket a REST client has to remember to add; the schema's type language lets a model compose a correct query from training-time knowledge instead of a discovery-then-dispatch chain; and the same join operation costs differently depending on who performs it — expensive and brittle in the agent's inference loop, cheap and deterministic in the router.

I ran two experiments. Phase 1 pointed MCP servers at GitHub's live API and asked the same
question through each — 24 runs, $0.53. Phase 2 built a synthetic three-service airline
backend, generated a REST surface and a federated GraphQL surface from a single field
definition, and swept four questions over how many records they cover — 240 runs, $51.16.
Everything ran on `claude-haiku-4-5` at temperature 0, through Goose, behind a logging reverse
proxy that recorded the raw Anthropic `usage` object for every model call.

*The per-cell tables, the mechanism, the scored pre-registration
and the caveats in full are in [`FINDINGS.md`](https://github.com/apollographql/graphql-mcp-benchmarks/blob/main/FINDINGS.md); the generated reports are in
[`results/`](https://github.com/apollographql/graphql-mcp-benchmarks/blob/main/results); how to run any of it is in [`README.md`](https://github.com/apollographql/graphql-mcp-benchmarks/blob/main/README.md).*

---

## What ran

### Phase 1 — GitHub's live API

Four conditions, two tasks, three replicates each. The narrow question: on a real API that ships
both a REST interface and a GraphQL interface, what does the same request cost through each?

| cell | protocol | packaging | server | tools |
| --- | --- | --- | --- | --- |
| `A1` | REST | every toolset the server ships | `github-mcp-server`, Docker, `--read-only` | 54 |
| `A2` | REST | `--toolsets repos,issues,pull_requests` | same binary | 22 |
| `B` | GraphQL | schema discovery then execute | `apollo-mcp-server` v1.14.0, dynamic | 4 |
| `B2` | GraphQL | schema discovery then execute | `servers/rover_schema_mcp.py` | 3 |

```
flowchart LR
    A1["A1 · 54 tools"] --> GMCP
    A2["A2 · 22 tools"] --> GMCP
    GMCP["github-mcp-server<br/>Docker · stdio · --read-only"]
    GMCP --> RST["api.github.com<br/>REST · no batching, no field selection"]

    B["B · 4 tools"] --> AMCP["apollo-mcp-server v1.14.0<br/>dynamic mode"]
    B2["B2 · 3 tools"] --> RMCP["servers/rover_schema_mcp.py"]
    SDL[("GitHub SDL<br/>fetched with rover at setup")] --> AMCP
    SDL --> RMCP
    AMCP --> GQL["api.github.com/graphql"]
    RMCP --> GQL
```

 Loading

The two REST conditions differ only in how much of it is switched on. The two
GraphQL conditions are similar as well, one uses an open-source GraphQL MCP server and another uses
a thin wrapper over an existing schema search and description utility; both hand the agent a
query language rather than pre-built operations. `T1` asks for five pull
requests and their changed files, which is the N+1 case; `T2` asks for one, and is a
single-entity control that carries no protocol claim.

Phase 1's strength is that nothing about it is synthetic. Its limit is that it's measuring both
the protocol and the implementation details of GitHub's production services.
Phase 2 fills this gap.

### Phase 2 — a contrived backend

Eight condition cells, reported as eight rows and never averaged together. Each is one MCP server
pointed at one of two surfaces over the same three services.

| cell | protocol | packaging | server | tools |
| --- | --- | --- | --- | --- |
| `M-R1-fat` | REST | one tool per endpoint, full payloads | `openapi_mcp.py --mode tools` | 9 |
| `M-R1-lean` | REST | same, honoring `?fields=` | `openapi_mcp.py --mode tools` | 9 |
| `M-R2-fat` | REST | generic tools over the OpenAPI spec | `openapi_mcp.py --mode discovery` | 3 |
| `M-R2-lean` | REST | same, honoring `?fields=` | `openapi_mcp.py --mode discovery` | 3 |
| `M-R3-fat` | REST | one generic HTTP tool, no spec at all | `openapi_mcp.py --mode bare` | 1 |
| `M-G1` | GraphQL | schema discovery then execute | `supergraph_mcp.py` | 3 |
| `M-G2` | GraphQL | persisted operations, one tool each | `apollo-mcp-server` v1.14.0 | 7 |
| `M-G3` | GraphQL | schema discovery then execute | `apollo-mcp-server` v1.14.0, dynamic | 3 |

The server column is what makes the axes separable. `M-R1`, `M-R2` and `M-R3` are one file in
three modes, so the REST axis varies packaging and nothing else; `M-G2` and `M-G3` are one binary
in two modes, so the GraphQL axis does too. `M-G3` closes the square — same implementation as
`M-G2` with different packaging, same packaging as `M-G1` with a different implementation. `M-G1`
is a control I wrote, not a product anyone can install, and it is reported alongside the
shipping equivalent rather than in place of it.

Three services: flight scheduling, fleet maintenance, crew personnel. Each are modeled after an airline
operations stack because it gives a natural three-way join: a flight is scheduled by one
service, flown by an aircraft owned by another, and crewed by people belonging to a third. Both
surfaces are generated from one field declaration and read the same records through the same
repository, so the implementation won't have hidden bias toward either protocol.

```
flowchart LR
    ENT["services/src/entities/*.ts<br/>one field declaration per service"]
    ENT --> GEN["codegen"]
    GEN --> SDL["generated/*/schema.graphql"]
    GEN --> OAS["generated/*/openapi.json"]

    ENT -.-> PAR{{"parity.test.ts<br/>the fairness gate"}}
    SDL -.-> PAR
    OAS -.-> PAR

    FIX[("one repository<br/>hash-pinned fixtures")]

    subgraph GQL ["GraphQL surface"]
        direction LR
        SUB["3 subgraphs · Apollo Server v5<br/>per-request DataLoaders"] --> ROUTER["Apollo Router v2.17.0<br/>:5000"]
    end

    subgraph RST ["REST surface"]
        direction LR
        REST["3 Node HTTP services<br/>GET /v2/... · 9 endpoints"]
    end

    SDL --> SUB
    OAS --> REST
    FIX --> SUB
    FIX --> REST

    ROUTER --> GC["the 3 GraphQL conditions"]
    REST --> RC["the 5 REST conditions"]
```

 Loading

`parity.test.ts` is the fairness gate, and it enforces that every canonical field must be reachable
on both REST and GraphQL, and REST may carry extra keys only when they are declared redundant and
derived from a
canonical field. Extra bytes are permitted but not extra information since that's the whole point; the
extra bytes are what the study measures.

REST was the steelman. I gave it an OpenAPI document generated from the implementation so it
can never be stale or partial. It contains nine endpoints across three services with one naming convention,
one envelope and one pagination scheme; batch-by-id on every collection; and a `?fields=`
sparse-fieldset bracket. This is an extremely generous setup IMO. Payloads are deliberately bloated
in ways that production APIs typically
are. This includes envelope wrappers, code/label twins, and denormalized nested objects. A flight comes back
with 46 fields under the `fat` profile. Cross-service expansion is the one thing REST is not
allowed: a service may link to another service's resource but never inline it, because that is
precisely the constraint that GraphQL Federation exists to solve.

### The phase-2 tasks

Four questions, swept over how many records they cover: M1 one service and batchable
(REST's best case), M2 one record across three services, M3 M2 swept over *N*, M4 the
list in one service and the predicate in another. All ten instances are quoted verbatim in
[`tasks/tasks.yaml`](https://github.com/apollographql/graphql-mcp-benchmarks/blob/main/tasks/tasks.yaml), which is the only place the wording lives. M3, the
cross-service join, reads:

```
For each of these flights — {{ids}} — determine whether every assigned
pilot (the captain and the first officer) holds a type rating for that
flight's aircraft model which is still current as of {{as_of}}. Report one
line per flight: the flight id, then yes or no. Cover all {{n}} flights.
```

### How a run was measured

Nothing is read out of the agent's own logs. Goose points at a logging reverse proxy that
forwards every model call verbatim and tees the response stream, so the numbers are the raw
`usage` object the API returned rather than a harness's re-reporting of it.

```
flowchart LR
    RCP["recipes/recipe_m_*.yaml<br/>byte-identical instruction block"] --> GOOSE

    GOOSE["Goose · temperature 0<br/>claude-haiku-4-5"]
    GOOSE -->|"stdio · tool calls"| MCP["the condition's MCP server"]
    MCP --> STACK[("the backend above")]

    GOOSE -->|"HTTP · ANTHROPIC_HOST"| FWD

    subgraph PX ["proxy/anthropic_logging_proxy.py"]
        FWD["forward verbatim<br/>headers + body unchanged"]
        TEE["tee the SSE stream"]
    end

    FWD --> API["api.anthropic.com"]
    API --> TEE
    TEE -->|"stream unchanged"| GOOSE

    TEE --> PJ[("proxy.jsonl<br/>raw usage object, one line per call")]
    TEE --> TIO[("tool_io.jsonl<br/>tool arguments + result bodies<br/>keyed by tool_use_id")]

    PJ --> PARSE["parse_logs.py · grade.py"]
    TIO --> PARSE
    PARSE --> OUT["results/** · figures/**"]
```

 Loading

The proxy is deliberately naive: it records what crossed the wire and decides nothing, because it
is the one component whose correctness underpins every published number. The sidecar exists
because the headline metric needs to know both the size and content of a payload.

That metric is "pass-through tokens:" payload that entered the agent's context and whose values
never appear in its answer. Put another way, it's the data the agent carried, paid for
on every subsequent call, and didn't use. In other words: waste. All five phase-2 recipes carry a byte-identical
instruction block that names no tool and suggests no strategy, and the runner refuses to start if
they drift.

---

## The result

GraphQL-over-MCP outperforms REST-over-MCP across all tasks.

[Every GraphQL condition carried less waste than every REST condition](https://github.com/apollographql/graphql-mcp-benchmarks/blob/main/figures/fig1-arm-separation.png)

All three GraphQL conditions place above all five REST conditions. The worst GraphQL condition carries
2.5× less than the best REST condition.

Best GraphQL cell against best REST cell, instance by instance:

| task | best REST | best GraphQL | cost ratio | token ratio |
| --- | --- | --- | --- | --- |
| M1 @ 1 | $0.0081 | $0.0046 | 1.76× | 15.71× |
| M1 @ 5 | $0.0145 | $0.0058 | 2.50× | 10.73× |
| M1 @ 20 | $0.0159 | $0.0111 | 1.43× | 1.18× |
| M1 @ 50 | $0.0251 | $0.0202 | 1.24× | 1.93× |
| M2 @ 1 | $0.0221 | $0.0068 | 3.25× | 3.55× |
| M3 @ 5 | $0.0858 | $0.0193 | 4.45× | 4.04× |
| M3 @ 20 | $0.2214 | $0.0480 | 4.61× | 2.89× |
| M3 @ 50 | $0.4765 | $0.0677 | 7.04× | 11.88× |
| M4 @ 20 | $0.0733 | $0.0233 | 3.15× | 4.95× |
| M4 @ 50 | $0.1287 | $0.0388 | 3.32× | 9.06× |

Median cell: 3.19× on cost, 4.49× on tokens. Ten out of ten.

There is no single multiple here, and the margin is not monotone in N. It collapses as the
batchable single-service question grows and widens on the cross-service joins:

[The margin tracks the shape of the question, not its size](https://github.com/a
jdauriemma20
🟧 echo.github ⭐Reports 264 runs using Claude Haiku 4.5 through Goose, with median best-cell REST-to-GraphQL ratios of 3.19× on cost and 4.49× on tokens in Apollo GraphQL——

Interpretation history

Decision trace