2026-10-11 16:37 UTC

MCPJam claims its released testing platform evaluates how external AI clients use an MCP server and gates releases on repeated tool-choice, argument, and goal-completion checks, potentially catching integration regressions that server conformance tests miss.

state: watchingheat: lowuncertainty: highconvergesscott: mediummcp agent-evaluation tool-testingMCPJamPrathmesh

What is this?

MCPJam is an open-source (Apache 2.0) testing and evaluation platform for MCP server developers, led by CEO Prathmesh: an npx-run Inspector plus CLI, SDK, conformance checks, and a hosted app that tests how external AI clients (16 client configurations, 170+ models) use your server rather than an agent you control — scoring tool choice, arguments, and goal completion across repeated runs, strictly pre-production, with CI wiring that gates PRs on eval pass-rate thresholds alongside conformance checks. The supplied snippets confirm a real, actively maintained surface: 2,000+ GitHub stars per a third-party roundup, monthly 'MCP client changelog' posts tracking host updates and what to test, MCP v2 stateless-migration coverage, and a 2026 testing-tools roundup listing several MCP eval platforms — weak corroboration that a testing category is forming. Advertised effectiveness is not demonstrated in the supplied material: the '98% suite health' figures are MCPJam's own demo data, adoption numbers are first-party and unverified, and no independent source shows these evals catching regressions that conformance tests miss. One conflict worth flagging: an unrelated 'MCP-Jest' framework also bills itself as the first MCP server testing framework, so the 'first' positioning is contested.

Why it matters to Scott

MCPJam independently ships what Reflexive Agent Design argues: simulated agent personas walked across external clients with scored tool-choice/argument/goal traces, gated in CI exactly as Evaluation-Driven Development prescribes — a dated receipt for Scott's ebook thesis and a concrete trial candidate against his production mcp-ip-wiki map→page→source toolbelt. The re-grounding strengthens existence evidence (third-party 2,000-star roundup, a testing-tools roundup suggesting a category forming) but effectiveness is still advertised, not demonstrated, so relevance holds at medium rather than rising.
ip:framework.reflexive-agent-designip:source.reflexive-agent-design-ebookip:concept.evaluation-driven-developmentdev:project.mcp-ip-wikiradar:oqoqo-agent-interface-evalsradar:mcp-server-agent-usabilityradar:concept.agent-evaluation
queries asked of Scott's wikis
  • MCP server toolbelt testing and conformance
  • evaluation-driven development CI merge gates
  • reflexive agent design self-checking loops
  • tool choice and argument errors agent usability
  • maintaining agent harnesses across client and model drift
  • simulated user personas scored eval sessions

Measured heat

now 0 pts/hpeak 21 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 573h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-17 20:30 (minted)⭐ origin echo-reconstructedMCPJam offers cross-client MCP testing, simulated user journeys, local evaluation tooling, and CI checks, with an open-source Inspector, CLI
MCPJam on blog (echo) · attributed from hn.story.49745351 · published time unknown
—
09-17 19:23first on hacker news · published · lag ?Show HN: MCPJam - the first testing & evaluations platform for MCP servers
prathmeshmcp
—
09-17 19:23amplified on hacker news 👑hn.story.49745351
prathmeshmcp
peak 13 · 9 comments · 79% of case engagement
10-05 13:00amplified on hacker newshn.story.49964286
pwizard234
peak 2 · 0 comments · 7% of case engagement
10-06 09:12amplified on hacker newshn.story.49976078
pwizard234
peak 4 · 0 comments · 14% of case engagement
09-17 20:21our radar first saw it · lag ?discovery anchor: hn.story.49745351—
pace: p55 vs 1032 stories at the 336h mark (now 573h old) — ahead of aws-project-spend-limits (1.1x), behind all-your-agents-session-monitor (0.9x)

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: MCPJam - the first testing & evaluations platform for MCP servers
Retrieved article excerpt

Open article · Retrieved 2026-09-17T20:23:05.628386+00:00

Product

[Inspector](https://www.mcpjam.com/features/inspector)[OAuth & EMA debugger](https://www.mcpjam.com/features/oauth-debugger)[Cross-client testing](https://www.mcpjam.com/features/cross-client-testing)[CI/CD actions](https://www.mcpjam.com/features/ci-cd-actions)[See all features](https://www.mcpjam.com/features)[CLI](https://www.mcpjam.com/features/cli)[SDK](https://www.mcpjam.com/features/sdk)[Compare client capabilities](https://www.mcpjam.com/clients)[MCPJam vs. other tools](https://www.mcpjam.com/compare)

Solutions

[For developers](https://www.mcpjam.com/for/developers)[For product managers](https://www.mcpjam.com/for/product-managers)[For engineering managers](https://www.mcpjam.com/for/engineering-managers)[For AI & platform leads](https://www.mcpjam.com/for/ai-platform-leads)[For enterprise](https://www.mcpjam.com/for/enterprise)[For marketers](https://www.mcpjam.com/for/marketers)[For MCP clients](https://www.mcpjam.com/for/mcp-clients)

[Pricing](https://www.mcpjam.com/pricing)[Blog](https://www.mcpjam.com/blog)[Docs](https://docs.mcpjam.com)

[Start testing](https://app.mcpjam.com)

# Test your MCP server in every major AI client.

From your first prompt to a continuous gate on every release, MCPJam shows what breaks across every AI client, and how to fix it.

Test in 16+ major AI clients

[Claude](https://www.mcpjam.com/clients/claude "Claude")[ChatGPT](https://www.mcpjam.com/clients/chatgpt "ChatGPT")[Cursor](https://www.mcpjam.com/clients/cursor "Cursor")[Copilot](https://www.mcpjam.com/clients/copilot "Copilot")[VS Code](https://www.mcpjam.com/clients/vscode "VS Code")[Claude Code](https://www.mcpjam.com/clients/claude-code "Claude Code")[Codex](https://www.mcpjam.com/clients/codex "Codex")[Perplexity](https://www.mcpjam.com/clients/perplexity "Perplexity")[+8](https://www.mcpjam.com/clients)

[MCPJam - The testing & evaluations platform for MCP servers | Product Hunt](https://www.producthunt.com/products/mcpjam-inspector?embed=true&utm_source=badge-featured&utm_medium=badge&utm_campaign=badge-mcpjam)

Get your agent set up with MCPJam

Copy setup prompt

Launch locally or in web

npx @mcpjam/inspector@latest

[Open web app](https://app.mcpjam.com)

MCPJam

MCPJam

PlaygroundEvaluateSwarmsUser Testing

Fillen-USSan FranciscoStrictClient contextHost capabilitiesClear chat

ChatGPTs-2291-c3

ChatTraceRaw

1. Show page views and signups for the last 7 days.
2. show\_analytics Complete

   PULSE · MCP APPDemo data

   Page views

   24,800

   Chart typeLineBar

   Sep 8–14 · Last 7 days

   02.5k5kSep 8: 2800Sep 9: 3400Sep 10: 3100Sep 11: 4200Sep 12: 3600Sep 13: 3300Sep 14: 4400Sep 8Sep 10Sep 12Sep 14

Claudes-2291-a1

ChatTraceRaw

1. Show page views and signups for the last 7 days.
2. show\_analytics Complete

   PULSE · MCP APPDemo data

   Signups

   386

   Chart typeLineBar

   Sep 8–14 · Last 7 days

   04080Sep 8: 38Sep 9: 62Sep 10: 45Sep 11: 71Sep 12: 54Sep 13: 48Sep 14: 68Sep 8Sep 10Sep 12Sep 14

Show page views and signups for the last 7 days.Refund the duplicate charge on order A-2291Which charges made up Tuesday's payout?

Ask something… Use Slash “/” commands for Skills & MCP prompts

ChatGPTClaudeCursor

## Trusted by 106,000+ developers and 280+ enterprises

- IBM
- Scalekit
- Apollo

[> “We use MCPJam every day. It has become essential for testing MCP servers locally.”

Jau Chan · Senior Software Engineer, Asana

See how they skip the 2-minute deploy](https://www.mcpjam.com/blog/asana-mcpjam)[> “If I run out of MCPJam credits, my day stops. I don't know how I would work without it.”

Atlas Wegman · Staff iOS Engineer, Jeppesen ForeFlight

See how they iterate 10x faster](https://www.mcpjam.com/blog/foreflight-mcpjam)

Inspector

## See the same prompt land in multiple clients.

Call a tool, read the trace, and compare how your product appears in different clients.

Fillen-USSan FranciscoStrictClient contextHost capabilitiesClear chat

ChatGPTs-2291-c3

ChatTraceRaw

Claudes-2291-a1

ChatTraceRaw

Show page views and signups for the last 7 days.Refund the duplicate charge on order A-2291Which charges made up Tuesday's payout?

Ask something… Use Slash “/” commands for Skills & MCP prompts

ChatGPTClaudeCursor

Swarms

## Test the journeys your users actually take.

Agent personas run them across every client. Every session is captured and scored.

1. 1Describe
2. 2Confirm details
3. 3Run swarm

Leave

### Create a swarm of your users

Simulated users run through your server so you can see what breaks.

Choose the clients and servers your users will interact with. \*

This will shape the swarm's behavior. You will be able to compare behavior across clients side by side.

ChatGPTClaudeCursor+ Add client

acme-payments· v2.15.0-rc.1 · 5 tools

Describe your users and their behavior. \*

We build the user goals from your input. Your inputs are not final. You can edit personas and goals on the next screen.

Support agents refunding duplicate charges, finance ops reconciling Tuesday's payouts, and developers wiring a failed-payment webhook.

Add existing persona

Generated personas

- MCMaya Chen · Support agentFind the duplicate charge on order A-2291 and refund it.8 × 3 clients
- JHJordan Hale · Finance opsMatch payout po\_7Hd2 to its charges and flag anything missing.8 × 3 clients
- PSPriya Shah · DeveloperCreate an invoice, fail its payment, and confirm the webhook fires.8 × 3 clients

CancelContinueRun swarm

User Testing

## Then invite real people to try the same ones.

Share a sandboxed chatbox. Measure user sentiment and uncover usability gaps.

Refund & reconcile · beta

Findings2InsightsSessions15

EditOpen previewShare

#### Access settings

Anyone with the link (guests included)

#### Has access

- MMara Okafor (you)[[email protected]](https://www.mcpjam.com/cdn-cgi/l/email-protection)Owner
- Dana Reyes[[email protected]](https://www.mcpjam.com/cdn-cgi/l/email-protection)Tester
- Luis Moreno[[email protected]](https://www.mcpjam.com/cdn-cgi/l/email-protection)Tester
- Sam Kaur[[email protected]](https://www.mcpjam.com/cdn-cgi/l/email-protection)Tester

#### Ratings

Per-turn ratings

starsthumbs

Preview

ChatGPTMCPJamCopy link

Refund the duplicate charge on order A-2291

list\_charges· refund\_charge

I found two charges for order A-2291: ch\_3RkA9 for $129.00 and ch\_3RkQ9 for $129.00, 41 seconds apart. I refunded ch\_3RkQ9 as a duplicate. The customer will see it in 5 to 10 business days.

Rate this response

Send

Message…

AI can make mistakes. Please double-check responses.

Previewing as you, signed in. A tester opening this link gets the guest experience, which can differ.

Per-turn rating · zoomed

How did the last few turns go?

Picked the right tool. Could've grouped by service.

Send →

Evaluate

## Set the quality bar for every client.

Lock in the behaviors you want always working. The next run tells you if they held, and in which client.

### Evaluate

OverviewSuites

SetupRun

Suite health

Payments operations

98%average across 12 runs

acme-paymentsPayments operations

100%95%

9/29/39/49/59/69/79/89/99/109/119/129/13

Runs

PlatformSuiteClientServer

| Run | Suite | Client | Result | Rate | Platform | Date | Latency | Tokens | Calls |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| #128 | Payments operations |  | Failed | 96%69/72 | GitHub | Sep 13, 2:00 AM | 18.4s | 16.2k | 44 |
| #127 | Payments operations |  | Passed | 98%71/72 | CLI | Sep 12, 2:00 AM | 14.2s | 12.1k | 41 |
| #126 | Treasury reconciliation |  | Passed | 100%48/48 | SDK | Sep 12, 1:14 AM | 16.8s | 15.6k | 38 |
| #125 | Billing exception handling |  | Failed | 94%45/48 | CLI | Sep 11, 11:42 PM | 21.1s | 18.9k | 47 |
| #124 | Payments operations |  | Passed | 99%71/72 | GitHub | Sep 11, 2:00 AM | 11.3s | 9.4k | 29 |
| #123 | Dispute resolution |  | Passed | 100%36/36 | SDK | Sep 10, 6:18 PM | 15.5s | 11.8k | 33 |
| #122 | Treasury reconciliation |  | Passed | 97%70/72 | GitHub | Sep 10, 2:00 AM | 17.9s | 14.7k | 40 |

CI/CD

## Block the ones that fail before they merge.

Run the same suite on every pull request. A failing check blocks the merge.

acme/acme-paymentsPrivate

CodeIssuesPull requestsActions

### list\_charges: return full charge objects #482

Open**m.okafor** wants to merge 1 commit into `release/2.15`

ConversationCommits 1Checks 5Files changed

Some checks were not successful

1 failing and 4 successful checks

- BuildSuccessfulRequired
- Unit testsSuccessfulRequired
- MCPJam / Protocol conformanceSuccessfulRequired
- MCPJam / OAuth & securitySuccessfulRequired
- MCPJam / Evals89% pass rate · 95% requiredRequired

Merging is blocked

MCPJam Evals must pass before this pull request can be merged.

Merge pull request

Questions

## Before you start jamming.

### Do I have to keep up with every AI client myself?

No. MCPJam maintains each client's current behavior so your tests reflect what users experience today.

### How do you test something non-deterministic?

With evaluation, not brittle assertions. MCPJam scores tool choice, arguments, and goal completion across repeated runs, then shows why the score changed.

### How is this different from Datadog, Braintrust, or LangSmith?

Those products evaluate an agent you control. MCPJam tests how external AI clients use your MCP server.

### How do you handle security and my data?

MCPJam is pre-production and never sees live traffic. The Inspector runs locally or in your CI; enterprise controls include SSO, audit logs, and a DPA.

### Is it open source?

Yes. The Inspector, CLI, SDK, local evals, and conformance checks are open source and free.

Ship knowing it works, for every user, in every client.

[Start testing](https://app.mcpjam.com)[Book a demo](https://www.mcpjam.com/contact)
prathmeshmcp139
🟧 echo.blog ⭐MCPJam offers cross-client MCP testing, simulated user journeys, local evaluation tooling, and CI checks, with an open-source Inspector, CLIMCPJam——
🟧 hnShow HN: Mcpward – black-box contract and security testing for MCP serverspwizard23420
🟧 hnShow HN: Mcpward – contract and security testing for MCP servers in CIpwizard23440

Interpretation history

Decision trace