2026-10-11 16:37 UTC

Armature claims its published coding-agent experiments show substantial differences in third-party service selection across agents and repository contexts, making agent choice and harness interaction design consequential controls on generated software dependencies.

state: watchingheat: lowuncertainty: highconvergesscott: highcoding-agents agent-harnesses coding-agent-tool-selection agent-evaluationArmature

What is this?

Armature is a YC-backed startup in the 'agent discoverability' business: per its YC page, it runs real coding agents (Claude Code, Codex, Cursor) inside a panel of repositories, measures which third-party services they install, and consults for vendors on the docs/SDK/content changes that get them picked. In September 2026 it published a study of 16,893 such sessions across 75 repositories, headlined by the finding that the three agents pick the same tool in only ~42% of cases, alongside public leaderboards framed around agent priors, expressed needs, and repository context. In October 2026 the same team launched Agent.reviews, a platform where AI agents read and write tool reviews, claiming 50k+ sessions measured. The web record confirms the company, the study's existence and framing, and its explicitly commercial purpose ('part of our broader work on how to influence coding agents choices'); adjacent arXiv work independently establishes that harness structure materially changes same-model agent outcomes and that coding-agent supply chains carry dependency-resolution exposures, but nothing in the snippets shows independent replication, trace audit, or third-party verification of Armature's specific selection findings — nor do the snippets confirm the YC batch, the 50k+ figure, or the co-founder's identity.

Why it matters to Scott

A YC-backed team has independently arrived, with dated public measurements, at Scott's Model-Plus-Harness Benchmark Unit extended into a domain his canon gestures at but doesn't measure — harness and agent choice as the hidden selector of generated software's dependencies (Sovereign Software Assurance's concern) — and Agent.reviews is a live instance of the agent-written-review trust surface his Agent Addressability framework already describes, built by a Gemini-orchestrated/Gemini-judged pipeline that exhibits his Correlated Checkers Pitfall and open to exactly the vendor-content manipulation his taint-tracking work governs. It rises to high because it is actionable rather than confirmatory: service selection is a concrete outcome his trace-backed agent-comparison fixtures could cheaply replicate or refute, harness choice becomes a new input to his dependency-assurance checks, and the Agent.reviews launch is what tips the case — an early agentic-web discoverability institution whose claims are entirely single-source and commercially interested, making independent replication through his own harnesses the natural next move and the manipulation-surface reading a publishing opportunity.
ip:concept.model-plus-harness-benchmark-unitip:framework.agent-addressabilityip:framework.sovereign-software-assurancedev:concept.trace-backed-agent-comparisonip:concept.correlated-checkers-pitfallip:concept.taint-trackingradar:concept.coding-agent-harnessesradar:concept.agent-harnessesradar:frontierharness-17x-cost-variationradar:coding-assistant-supply-chain-trustradar:manufactured-ai-recommendation-sourcesradar:notion-mcp-undisclosed-upsellradar:concept.llm-judges
queries asked of Scott's wikis
  • model-plus-harness benchmark unit — does harness choice change downstream outcomes
  • generated code dependency supply chain assurance checks
  • LLM-as-judge reliability in agent trace evaluation
  • agentic web discoverability — SEO-for-agents, agent-written reviews as trust surface
  • vendor documentation as agent manipulation / prompt-injection surface
  • Claude Code vs Codex harness interaction design differences

Measured heat

now 0 pts/hpeak 8 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 719h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-11 17:24 (minted)⭐ origin echo-reconstructedArmature reports 16,893 experimental sessions and a first published subset of 5,292 valid sessions on 51 synthetic codebases; all three agen
The Armature team on blog (echo) · attributed from reddit.post.1wdlkwm · published time unknown
—
09-11 16:45first on r/ClaudeAI · published · lag ?Which tools does Claude Code choose and how does it compare to other agents? We measured 17k runs to find out
Prize_Invite_3244
—
10-07 16:59first on hacker news · published · lag ?Show HN: Agent.reviews – where AI agents read and write reviews on tools
screm
—
09-11 16:45amplified on r/ClaudeAIreddit.post.1wdlkwm
Prize_Invite_3244
peak 2 · 10 comments · 5% of case engagement
09-18 15:13amplified on r/ClaudeAIreddit.post.1wjsxm4
Prize_Invite_3244
peak 1 · 1 comments · 1% of case engagement
10-07 16:59amplified on hacker news 👑hn.story.49995539
screm
peak 72 · 49 comments · 94% of case engagement
09-11 17:21our radar first saw it · lag ?discovery anchor: reddit.post.1wdlkwm—
pace: p50 vs 1032 stories at the 336h mark (now 719h old) — ahead of agentgit-accountless-agent-handoffs (1.1x), behind anthropic-pentagon-blacklist-ruling (0.9x)

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditWhich tools does Claude Code choose and how does it compare to other agents? We measured 17k runs to find out
ClaudeAI
Retrieved article excerpt

Open article · Retrieved 2026-09-11T17:23:31.611686+00:00

How did we run all these experiments concretely? Our panel of repositories We started by running an analysis over thousands of public GitHub repositories from which we extracted statistics about programming languages & frameworks, third-party services, deployment platform, team sizes, and codebase age. Since Tech startups are more likely to have open-source repositories than large enterprises, and stacks are likely very different we then unbiased our statistics based on publicly available data and reached our ideal panel distribution. We then staffed various coding agents to create real-world repositories to match these exact requirements. Finally, we generated variants in which we removed parts of the codebases and with them, entire third-party service implementations so we could run proper unbiased experiments. We landed on 75 repositories, in 10 languages, all using fake company names, fake git histories, fake API keys and real lockfiles checked against package manager registries like npm. Real-world tasks Each experiment is a real task to be performed inside a repository, asked by one of the following 4 profiles: Vibe-coder: only describes symptoms and ideal state, rarely the tool category name Junior engineer: usually mentions the desired state and the category name Senior engineer: is more precise about requirements and things to avoid Engineer at a large enterprise: details specific constraints, compliance, procurement, etc. Prompts are generally simple and direct and slightly tailored to each experiment (taking into account on the repository and the persona) but in 20-25% of the cases we tested adding specific mentions to the prompts like costs or usage volume to test their impact on the final output. We ended up with 1,163 variations like this one: “Now I need that each invoice that we generate gets sent to the user’s email address with a nice message, find the best solution and implement it”. Runner Each experiment is run in a dedicated ephemeral sandbox. We verified that the choice of the sandbox didn’t impact the conclusions but just to be safe we decided to rotate between 3 different sandbox providers (namely E2B, Blaxel and Daytona). A “simulated human” in the loop Since real-world conversations are rarely just one prompt and an agent working continuously on its goal with no interruption, we decided to use a “simulated human” in the loop. We achieved this using an orchestrator, played by Gemini 3.7 Flash . This allowed us to play more realistic scenarios where the agent would be first asked to analyze the codebase and recommend the best solution. At this stage the simulated human would always go with the top 1 solution or ask the coding agent to choose the best one and implement it. But we noticed that asking at the beginning to implement without returning any question would bias the agent towards building everything in-house as it was not able to ask authorization to pick a specific third-party solution. Adding this “human” in the loop reduced the leaders & cloud platform-native solutions dominance towards a more realistic picture. For example in the object storage experiment, Cloudflare R2 started winning in sessions in which the agent would always use Amazon S3 before. Our judge Another instance of Gemini 3.7 Flash was used to analyze the sessions. Its role is twofold: Assess if a session is valid regarding a list of criterias, e.g., the choice wasn’t biased by a repository that already “pre-chose” the provider; a solution was actually chosen (for observability it would reject OpenTelemetry alone if not coupled with a platform). Identify each player that was mentioned, and the final winner (looking at the conversation and the actual code diffs). So what did we learn? Out of these 16,893 runs, we started by keeping 5,292 sessions on 51 codebases and 18 sectors that we considered valid and ready to be published. This doesn’t mean we threw the 10k+ others to the bin and may share them in a second wave. On this first wave, we only extracted a fraction of all the learnings that are still buried in the traces and will continue digging to share what surprised us and what’s of interest to vendors and developers. But from today, all these traces are public so you can do the same. Below are 5 first observations we found interesting. Different coding agents use different sources and they end up disagreeing. Cursor bases its decision on the web in 2/3 of the sessions. Codex almost always uses web search (94% of sessions) but in 9 queries out of 10 it uses operators like site: to focus on trusted domains or dive on a specific solution (like in site:auth0.com password reset MFA social connections for example) Claude Code relies primarily on its priors and searches the web only in ~30% of the cases. But when it does, it browses 3x more pages than Codex. In more recent sectors such as sandboxes where its priors are weaker, it searched the web ~80% of the time. All three agents pick the same tool in only 42% of the cells: in the voice agents category for example, Claude Code picks Twilio while Codex picks OpenAI Realtime API (👀) and Cursor goes with Vapi. Claude Code builds in-house almost twice as much as Codex and Cursor (19% vs 10%) Repository context is key With the exact same ask on 4 repositories in 4 different programming languages, we got 4 different email provider winners: Resend wins on Typescript (55/89 runs), Sendgrid on Python (22/24), Postmark on Go (20/24) and Azure ACS on Java (22/23). While Vercel wins on Typescript repos (and naturally, even in 100% of the case when NextJS is used), it was never recommended on Python repos where Render dominated. Getting mentioned isn’t winning So many well-known players are mentioned in almost every conversation and are never picked. Of course, in the real world you’d expect a share of them to still win because of human involvement in the choice but some results are striking: In the payment service provider sector, Paypal is cited 139 times and never picked (Stripe won 124 of these 139 sessions). Same for Adyen mentioned 175 times and picked 3 times only. LangChain is the most cited framework with 194 mentions but was only picked 4 times (!). Netlify was mentioned 152 times and picked 6 times as the deployment platform. Supabase is the most mentioned database with 242 mentions and was still largely dominated by Neon. Additional features or details on vendors pages can flip choices Mailgun regularly lost against Postmark when agents read “1-day retention” on its free plan Supabase almost always lost because of too many unnecessary BaaS features (auth, storage, realtime) presented in a bundle pricing while agents were looking for a database only Out of our 5.3k sessions, 388 mentioned platform management overhead and 195 mentioned costs. In a significant of these cases, we noticed that this was more due to a way of presenting the information rather than an actual disqualifying datapoint. Some markets are outrageously dominated, some are very disputed Stripe won in 9 cases out of 10, losing only in specific EU-regulated cases where some players were more specialized (Paddle, Mollie). Neon won on 66% followed by cloud platforms native solution (Azure, AWS). For File storage Amazon S3 dominates with 45% followed by Azure and GCP with 20% each Resend and Postmark lead closely with respectively 35.6% and 27.4% of install rate. This is only the beginning of our experiments and we’ll keep publishing insights about how coding agents choose third-party services. We also plan to run brand new experiments so we’d like to know what are the questions you still have, don’t hesitate to reach out to us at [email protected] . Who wins in each sector? Why? To answer those burning questions, we are exposing all our results with our analyses, key learnings and entire traces in the leaderboard below!
Prize_Invite_3244210
🟧 echo.blog ⭐Armature reports 16,893 experimental sessions and a first published subset of 5,292 valid sessions on 51 synthetic codebases; all three agenThe Armature team——
🟠 redditI asked Claude Code to pick the "best" code review tool and it chose itself 76% of the time
ClaudeAI
Prize_Invite_324401
🟧 hnShow HN: Agent.reviews – where AI agents read and write reviews on toolsscrem7249

Interpretation history

Decision trace