2026-10-11 16:37 UTC

simonether's pre-registered, placebo-controlled trial of the 9 most-starred Claude Code skills reports only 2 beat a token-matched neutral placebo (planning-with-files does worse) and none beats no-skill on cost โ€” and whether the method spreads (the independent Sonnet Ponytail replication, further harness ports, skill-author disputes and responses) decides if the skills ecosystem shifts to measured validation, while a methodological rebuttal or fade closes it.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: highagent-evaluation claude-code-skills benchmark-methodologysimonetherDietrichGebert

What is this?

Anthropic's Claude Code supports 'skills' โ€” folder-based instruction packs (a SKILL.md plus optional scripts) injected into context when a task matches โ€” and the ecosystem has grown past 60,000 published skills, with GitHub stars driving curation and skills' standing token overhead already a mainstream complaint. Against that backdrop, simonether ran a pre-registered, drug-trial-style efficacy test of the 9 most-starred Claude Code skills: 450 trials comparing each skill against a token-matched neutral placebo (a same-length instruction pack that says nothing useful) and against no skill at all. Per the case's own evidence, only 2 of 9 beat placebo (both at Holm-adjusted p = 0.049), the Manus-style planning-with-files skill (~15.5k stars) performed worse than the pill, none beat no-skill on cost, and an independent replication (the Ponytail skill on Sonnet, 24 runs) found 23% less code but no token or time savings. The supplied web snippets do not directly surface the trial itself or its disputed authors โ€” they establish the surrounding ecosystem, its star-driven curation, and the standing concern about skills' context-token cost that the trial is stress-testing.

Why it matters to Scott

simonether's design is Scott's own Brief A/B Testing with a token-matched placebo arm, run to Falsifiability-Spine discipline (pre-registered skills, pre-written falsifiers, Holm-adjusted acceptance) โ€” the skills ecosystem independently arriving at his measured-validation methodology โ€” and the headline result (most-starred instruction packs don't beat same-length nothing; none beat no-skill on cost) is the Fat AGENTS.md anti-pattern confirmed at ecosystem scale with dated receipts. The planning-with-files skill doing worse than the pill both supports his less-prescriptive north-star prompting doctrine and forces a sharpening of his Markdown-OS claim โ€” injected prescriptive planning packs are not agent-written durable state โ€” making this a publishing opportunity plus a placebo-control method he can port into his own skill/brief harnesses.
ip:concept.fat-agents-md-anti-patterndev:concept.brief-ab-testingip:framework.falsifiability-spineip:concept.skills-and-workflowsdev:technology.claude-codeip:framework.markdown-osip:concept.north-star-promptingradar:agent-skill-bloat-gradingradar:concept.agent-skillsradar:concept.agent-evaluationradar:concept.claude-coderadar:concept.benchmark-integrityradar:concept.token-efficiencyradar:driftproof-skill-regression-testing
queries asked of Scott's wikis
  • Claude Code skills SKILL.md authoring patterns
  • agent memory persistent planning markdown files
  • agent evaluation harness statistical methodology
  • context token cost bloat CLAUDE.md overhead
  • pre-registered measured validation agent tooling claims
  • Ponytail skill token savings

Measured heat

now 0 pts/hpeak 15 pts/hcomments 0/hpeers p33momentum: steady3 platformsage 290h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-29 14:00โญ origin echo-reconstructedPlacebo-controlled test of the most-starred coding agent skills: '2 of 9 skills beat a same-length placebo, both at Holm-adjusted p = 0.049
simonether on github (echo) ยท attributed from hn.story.49979220, reddit.post.1wz4rfz, reddit.post.1wz4677
โ€”
10-06 14:33first on r/ClaudeAI ยท published ยท +168.6hI tested the Ponytail skill on Sonnet (24 runs): 23% less code, but tokens and time didn't drop
Sufficient-Storage87
โ€”
10-06 14:38first on hacker news ยท published ยท +168.7hShow HN: Popular Claude Code skills vs. a placebo: 2 beat it, 1 did worse
simonether
โ€”
10-06 14:33amplified on r/ClaudeAIreddit.post.1wz4677
Sufficient-Storage87
peak 0 ยท 3 comments ยท 4% of case engagement
10-06 14:38amplified on hacker newshn.story.49979220
simonether
peak 1 ยท 0 comments ยท 3% of case engagement
10-06 14:58amplified on r/ClaudeAI ๐Ÿ‘‘reddit.post.1wz4rfz
simon_ether
peak 24 ยท 18 comments ยท 61% of case engagement
10-07 05:49amplified on r/ClaudeAIreddit.post.1wzoq5a
loonpwn
peak 0 ยท 4 comments ยท 6% of case engagement
10-08 08:22amplified on r/ClaudeAIreddit.post.1x0lg54
maverick_man1111
peak 2 ยท 6 comments ยท 12% of case engagement
10-09 15:09amplified on r/ClaudeAIreddit.post.1x1nvh8
Sufficient-Storage87
peak 3 ยท 7 comments ยท 15% of case engagement
10-06 16:24our radar first saw it ยท +170.4hdiscovery anchor: hn.story.49979220โ€”

Evidence (7) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnShow HN: Popular Claude Code skills vs. a placebo: 2 beat it, 1 did worse
Retrieved article excerpt

Open article ยท Retrieved 2026-10-06T16:42:29.607607+00:00

Forest plot: cost ratio of each skill vs its same-length placebo with 95% confidence intervals

# skill-placebo

**A placebo-controlled test of the most-starred skills for coding agents: does a skill do better than the same amount of neutral text?**

> Drug trials give the control group a sugar pill. We gave coding agents one: neutral instructions of the same length, installed the same way as each skill.

**2 of 9 skills beat a same-length placebo, both at Holm-adjusted p = 0.049 (1 worse, 6 no better).** Control: a same-length neutral placebo, installed like the skill.  
Secondary, against no skill: none of the 9 skills was measurably cheaper than running without a skill; a same-length placebo alone changed cost by +2% to +16%.  
Measured 2026-09-30 on 15 public tasks (SWE-bench Verified, Terminal-Bench 2.1, OpenThoughts-TBLite) with claude-opus-5-5 in Claude Code, 30 trials per arm, 450 trials in total. Method registered before the first run: [METHOD.md](https://github.com/simonether/skill-placebo/blob/main/METHOD.md). Per-trial records: [`results/`](https://github.com/simonether/skill-placebo/tree/main/results/); the agents' full logs: [release v0.1.0](https://github.com/simonether/skill-placebo/releases/tag/v0.1.0) (sha256 in `results/*/*/AGENT_LOGS.json`). Reproduce a row: `uvx skill-placebo run <owner/repo>`.

## Results: Claude Code (claude-opus-5-5)

| Skill | Cost vs placebo R [95% CI] | Change | Pass skill / placebo | D, pp [95% CI] | Verdict | n |
| --- | --- | --- | --- | --- | --- | --- |
| superpowers โ€  | 0.98 [0.91, 1.06] | โˆ’2% | 83% / 87% | โˆ’3 [โˆ’10, +0] | no better than placebo | 30/30 |
| mattpocock | 1.05 [0.99, 1.12] | +5% | 83% / 87% | โˆ’3 [โˆ’10, +0] | no better than placebo | 30/30 |
| karpathy | 1.09 [0.93, 1.23] | +9% | 83% / 87% | โˆ’3 [โˆ’13, +7] | no better than placebo | 30/30 |
| ponytail โ€  | 0.88 [0.76, 0.97] | โˆ’12% | 90% / 87% | +3 [โˆ’7, +13] | beats placebo | 30/30 |
| caveman | 0.99 [0.91, 1.06] | โˆ’1% | 90% / 83% | +7 [+0, +17] | no better than placebo | 30/30 |
| agent-skills | 0.95 [0.90, 0.99] | โˆ’5% | 83% / 83% | +0 [โˆ’10, +10] | beats placebo | 30/30 |
| i-have-adhd โ€  | 0.92 [0.86, 0.98] | โˆ’8% | 87% / 87% | +0 [โˆ’10, +10] | no better than placebo | 30/30 |
| planning-with-files | 1.05 [0.91, 1.19] | +5% | 80% / 100% | โˆ’20 [โˆ’37, โˆ’7] | worse than placebo | 30/30 |
| compound-engineering | 1.04 [0.91, 1.23] | +4% | 87% / 87% | +0 [โˆ’10, +10] | no better than placebo | 30/30 |

CIs are unadjusted; verdicts use Holm-adjusted p across 9 skills (i-have-adhd: cost CI excludes 1, Holm-adjusted p = 0.095).  
โ€  Scripted approval, not an author-documented mode ([METHOD.md 5.1](https://github.com/simonether/skill-placebo/blob/main/METHOD.md#51-scripted-approval-turn)): the scripted turn fired in 4 of 448 recorded trials (ponytail 1, i-have-adhd 2, placebo cc-3 1), never for superpowers.  
Corrected on 2026-10-06 ([METHOD.md, amendment 19](https://github.com/simonether/skill-placebo/blob/main/METHOD.md#16-amendments)): two agent timeouts whose tests passed afterwards now count as failed trials, as sections 5 and 7 require; planning-with-files moves from "no better" to "worse than placebo". Costs are unchanged.

R is the skill's mean cost divided by its placebo's; below 1 the skill is cheaper. D is the pass-rate
difference. Verdicts follow [METHOD.md 9.1](https://github.com/simonether/skill-placebo/blob/main/METHOD.md#91-verdict-per-skill-and-harness), with Holm
correction across the 9 skills. D stays out of the headline: with this many trials its 95% CI is about ยฑ10 points, too wide to rank skills by. This is v1: 30 trials per arm
([METHOD.md, amendments 16-17](https://github.com/simonether/skill-placebo/blob/main/METHOD.md)); more runs come as updates.

## Codex (gpt-6-sol): pilot only, secondary

| Skill | Cost vs placebo R [95% CI] | Change | Pass skill / placebo | D, pp [95% CI] | Verdict | n |
| --- | --- | --- | --- | --- | --- | --- |
| ponytail | 1.09 [0.90, 1.23] | +9% | 90% / 100% | โˆ’10 [โˆ’30, +0] | no better than placebo | 10/10 |
| agent-skills | 1.24 [1.10, 1.42] | +24% | 100% / 80% | +20 [+0, +60] | worse than placebo | 10/10 |
| compound-engineering | 1.42 [1.23, 1.64] | +42% | 100% / 100% | +0 [+0, +0] | worse than placebo | 10/10 |

These are the pilot's kill test on Codex: 3 skills on 5 tasks, 10 trials per arm.
The Codex main run did not take place for v1 ([METHOD.md, amendment 17](https://github.com/simonether/skill-placebo/blob/main/METHOD.md)); it may come as an update.

## What the READMEs claim, and what we measured

| Skill | Claimed (quote, source) | Their setup | Our closest measure | Measured |
| --- | --- | --- | --- | --- |
| superpowers | no numeric claim in README at [`8ca22db`](https://github.com/obra/superpowers/tree/8ca22dba9a94f28898bbce59f2537ff4d87c747d) |  |  | verdict only |
| mattpocock | no numeric claim in README at [`c55ee46`](https://github.com/mattpocock/skills/tree/c55ee46073ed923f86ce59a5eb3b6d895095d1b7) |  |  | verdict only |
| karpathy | no numeric claim in README at [`2c60614`](https://github.com/multica-ai/andrej-karpathy-skills/tree/2c606141936f1eeef17fa3043a72095b4765b9c2) |  |  | verdict only |
| ponytail | ~54% less code ([README.md:33](https://github.com/DietrichGebert/ponytail/blob/e3ba2aa6f1e6f0bc4d69eb09c9f0d0a93af56156/README.md#L33)) | mean over 12 feature tasks, Haiku 4.5, n=4, vs the same agent with no skill (README.md:34) | final diff size vs baseline: lines added + deleted + untracked, SWE-bench tasks only, trials after the measure was switched on (METHOD.md amendment 12a) | โˆ’13%, 95% CI [โˆ’27%, +12%]; n = 11 / 10 trials on 7 tasks (few trials: read with care) |
| ponytail | ~20% cheaper ([README.md:33](https://github.com/DietrichGebert/ponytail/blob/e3ba2aa6f1e6f0bc4d69eb09c9f0d0a93af56156/README.md#L33)) | mean over 12 feature tasks, Haiku 4.5, n=4, vs the same agent with no skill (README.md:34) | cost vs baseline (and vs placebo) | โˆ’1%, 95% CI [โˆ’19%, +13%]; n = 30 / 30 trials on 15 tasks |
| ponytail | ~27% faster ([README.md:33](https://github.com/DietrichGebert/ponytail/blob/e3ba2aa6f1e6f0bc4d69eb09c9f0d0a93af56156/README.md#L33)) | mean over 12 feature tasks, Haiku 4.5, n=4, vs the same agent with no skill (README.md:34) | agent wall time vs baseline | โˆ’11%, 95% CI [โˆ’30%, +13%]; n = 30 / 30 trials on 15 tasks |
| ponytail | -22% tokens ([README.md:85](https://github.com/DietrichGebert/ponytail/blob/e3ba2aa6f1e6f0bc4d69eb09c9f0d0a93af56156/README.md#L85)) | same benchmark; the -22% column is tokens | total tokens vs baseline | +4%, 95% CI [โˆ’11%, +18%]; n = 30 / 30 trials on 15 tasks |
| caveman | 65% fewer tokens (talking like a caveman: the agent's output) ([GitHub repository description](https://github.com/JuliusBrussee/caveman)) | repository description, read with gh api on 2026-09-30 | output tokens vs baseline (comparable to the claim) | โˆ’2%, 95% CI [โˆ’9%, +6%]; n = 30 / 30 trials on 15 tasks |
| caveman | 65% fewer tokens ([GitHub repository description](https://github.com/JuliusBrussee/caveman)) | repository description, read with gh api on 2026-09-30 | total tokens vs baseline (input and cache included; not the claim's metric) | +18%, 95% CI [+8%, +28%]; n = 30 / 30 trials on 15 tasks |
| caveman | 65% fewer tokens ([GitHub repository description](https://github.com/JuliusBrussee/caveman)) | repository description, read with gh api on 2026-09-30 | cost vs baseline (not the claim's metric) | +14%, 95% CI [+8%, +21%]; n = 30 / 30 trials on 15 tasks |
| caveman | -50% output tokens vs a terse control ([README.md:206](https://github.com/JuliusBrussee/caveman/blob/2fd153c67988e980fb0b2455c90832159a6a5a25/README.md#L206)) | ten dev questions, skill vs a plain 'Answer concisely.' control, claude-opus-4-6 | output tokens vs placebo (our placebo is neutral, not terse) | โˆ’5%, 95% CI [โˆ’14%, +3%] |
| agent-skills | no numeric claim in README at [`2686b62`](https://github.com/addyosmani/agent-skills/tree/2686b620fc1fed2e8f60c704839c766b8594c6b6) |  |  | verdict only |
| i-have-adhd | no numeric claim in README at [`839872f`](https://github.com/ayghri/i-have-adhd/tree/839872f9d1cd634fed642b4589ce7226199cc15f) |  |  | verdict only |
| planning-with-files | 96.7% assertion pass rate ([README.md:33](https://github.com/OthmanAdi/planning-with-files/blob/51c1caa27f9fefe259e45a7cc92fa79ee8787cd7/README.md#L33)) | v2.21.0 eval on claude-sonnet-4-6, 30 assertions of file-pattern fidelity, not task success (README.md:702) | not comparable: our pass is the task's own tests; pass rate vs baseline and placebo is shown | 80% vs 87% without the skill, D โˆ’7 [โˆ’20, +7] pp |
| planning-with-files | 13.3 โ†’ 5.0 turns after a context wipe ([README.md:75](https://github.com/OthmanAdi/planning-with-files/blob/51c1caa27f9fefe259e45a7cc92fa79ee8787cd7/README.md#L75)) | turns to resume after a context wipe, internal benchmark v1 | not measured: single-session tasks, no context wipe (METHOD.md section 13) | not measured by this design |
| compound-engineering | no numeric claim in README at [`e80c5c4`](https://github.com/EveryInc/compound-engineering-plugin/tree/e80c5c40440b90672d78f032f6dfaedc0daeb292) |  |  | verdict only |

The claims were measured by their authors on other tasks, models and baselines (the "their setup"
column), so a gap between the two columns does not mean the claim was wrong. It shows how much the
number changes in a different setup.

## How it works

- Every skill runs in three arms on the same tasks: no skill, placebo, skill.
- The placebo is neutral text sized to the skill's always-on token footprint (within ยฑ10%, measured),
  delivered through the same mechanism: plugin, hook or memory file ([METHOD.md 4.1](https://github.com/simonether/skill-placebo/blob/main/METHOD.md#41-placebo-construction)).
- Cost is recorded tokens times public list prices: an estimate by tokens, not an invoice. The 95% CIs come from a cluster bootstrap over tasks.
- Skills are installed as their authors document, at a pinned commit ([skills.lock.json](https://github.com/simonether/skill-placebo/blob/main/skills.lock.json)).
  Every trial is published, including the failed and interrupted ones.

## For skill authors

If your skill was installed wrong or its placebo is unfair to it, [open an issue](https://github.com/simonether/skill-placebo/issues/new?template=skill-install-dispute.yml): skill, commit, what is
wrong, how to check. Every trial is public, so you can point at the exact runs. A confirmed installation error
means your skill's arms are rerun in full, the result is updated, an amendment in METHOD.md records it, and the
issue is linked here ([METHOD.md, section 14](https://github.com/simonether/skill-placebo/blob/main/METHOD.md#14-fairness-to-skill-authors)).

## Reproduce

```
uvx skill-placebo run DietrichGebert/ponytail   # one skill: baseline, placebo, skill on the same tasks
```

## Limits

- One model in the main run: claude-opus-5-5 in Claude Code at medium effort, the harness default. Codex (gpt-6-sol)
  ran only in a small pilot. Other models can behave differently.
- The tasks are public and probably in the models' training data. That affects every arm equally, but
  absolute pass rates say little about new work.
- Most tasks sat at the ceiling for claude-opus-5-5, so the pass rate carries little information here and
  the headline is cost.
- Tasks take minutes. Skills that pay off over long sessions (memory, multi-day plans) are not measured.

## Port me

- Run a skill on OpenCode or Gemini CLI (`good first issue`)
- Propose a skill: open an issue with the repository
- Rerun on fresh tasks (a recent SWE-rebench slice)

## License

MIT, see [LICENSE](https://github.com/simonether/skill-placebo/blob/main/LICENSE).

## Badge for tested skills

`[![placebo-tested](https://raw.githubusercontent.com/simonether/skill-placebo/main/docs/badges/<skill>-cc.svg)](https://github.com/simonether/skill-placebo#results-claude-code-claude-opus-5-5)`

---

Made by Simon (@simonether) ยท I take AI-built apps from demo to pr
simonether10
๐ŸŸ  redditI gave the 9 most popular Claude Code skills a sugar pill. 2 beat it, 1 did worse than the pill.
ClaudeAI
simon_ether2418
๐ŸŸ  redditI tested the Ponytail skill on Sonnet (24 runs): 23% less code, but tokens and time didn't drop
ClaudeAI
Sufficient-Storage8703
๐ŸŸง echo.github โญPlacebo-controlled test of the most-starred coding agent skills: '2 of 9 skills beat a same-length placebo, both at Holm-adjusted p = 0.049 simonetherโ€”โ€”
๐ŸŸ  redditskilld.dev: an open-source skills registry with real skill output from Opus 5.5
ClaudeAI
loonpwn04
๐ŸŸ  redditIf you tested your Claude Code skill on Haiku 5.5 once, you don't have a result. You have a screenshot.
ClaudeAI
maverick_man111116
๐ŸŸ  redditI tested a popular Claude Code debugging skill against a same-length placebo. Claude never used it on its own, and forcing it didn't change the score
ClaudeAI
Sufficient-Storage8737

Interpretation history

Decision trace