The case concerns a reported dataset finding that 9.70% of published Claude Code skill artifacts fail to load, raising questions about validation and packaging in the coding-agent skill ecosystem. The supplied snippets describe skills as folders containing a `SKILL.md` plus optional scripts or templates, and distinguish load or activation failures from failures during execution; they also recommend testing skills in a fresh session. However, the search results do not independently surface the cited Zenodo dataset, its methodology, Kynth Studio’s role, or Anthropic’s response, so the one-in-ten figure remains unverified from the provided web evidence.
The reported failure rate independently supports Scott’s position that reusable agent skills are production artifacts that need repeatable evaluation and binding release gates, rather than informal Markdown dropped into a folder. If independently reproduced, it could sharpen his skills-and-workflows doctrine into concrete cold-start loading checks and packaging safeguards; for now, the dataset and methodology remain unverified in the supplied evidence.
ip:concept.skills-and-workflowsip:concept.evaluation-driven-developmentip:framework.12-factor-agents-frameworkradar:concept.agent-skillsradar:shared-agent-skills-standard
queries asked of Scott's wikis
- agent skill packaging and validation
- coding-agent harness extension reliability
- executable skills as software dependencies
- agent capability registries and package managers
- cold-start testing for agent instructions
- supply-chain safeguards for agent plugins
2026-08-25T20:35:13Z
The case has produced no independent runtime reproduction, disclosed validation advance, or Anthropic/ecosystem response after repeated checks. The packaging concern remains plausible, but the specific one-in-ten estimate has faded as an active developing signal.
2026-08-23T20:32:03Z
New discussion sharpens the unresolved methodological issue: static checks of frontmatter and directory shape may not measure whether a skill loads in a real Claude Code installation. This adds useful scrutiny but no reproduction, ecosystem response, or evidence supporting the one-in-ten prevalence estimate.
2026-08-23T14:30:35Z
The newly attached corpus-wide claim is another promotion of the SkillWorks census, not an independent reproduction, so it does not corroborate the roughly one-in-ten estimate. The practical packaging failure modes remain plausible, but prevalence and classification methodology still need outside validation or an Anthropic response.
2026-08-23T14:22:58Z
evidence attached: reddit.post.1vw78dt — Independent corpus-wide scoring reports that a large share of published Claude Code skills fail to load, materially corroborating the open packaging-reliability hypothesis.
2026-08-22T12:29:04Z
SkillWorks exposes per-skill scoring and makes the originating census more inspectable, but it is a product of the same organization rather than an independent reproduction. The 9.70% estimate therefore remains uncorroborated, with no Anthropic or ecosystem response yet.
2026-08-22T12:22:49Z
evidence attached: hn.story.49398601 — Independent SkillWorks scoring directly tests the reported Claude Code skill loading failure rate and could corroborate the open case.
2026-08-20T20:38:58Z
The transcript audit adds an independent practical signal that installed skills often go unused or undiscovered, broadening the reliability concern beyond packaging. It does not reproduce the census’s 9.70% load-failure claim or distinguish malfunction from lack of need, so the central estimate remains uncorroborated.
2026-08-20T16:24:01Z
evidence attached: reddit.post.1vtnzzb — The user's transcript-based measurement independently supports concerns about Claude Code skill discovery or activation failures.
2026-08-20T12:46:03Z
No independent reproduction, implementation response, or Anthropic acknowledgement has appeared; the one-in-ten failure claim remains a single-source census rather than a corroborated ecosystem reliability finding.
2026-08-20T12:42:15Z
grounded: converges/medium — The reported failure rate independently supports Scott’s position that reusable agent skills are production artifacts that need repeatable evaluation and bindin
2026-08-20T12:39:36Z
origin walked (codex/luna, conf 0.97): anchor hn.story.49373129 -> echo.other.090fd96f73 by Kynth Studios
2026-08-20T12:37:50Z
case created — The linked census presents a bounded, testable reliability claim about the published Claude Code skill ecosystem.