2026-10-11 17:10 UTC

Independent testing will determine whether instruction-bloated agent skills materially impair skill selection or task performance and whether automated grading can identify the harmful patterns.

state: watchingheat: mediumuncertainty: highknownscott: mediumagent-skills coding-agents context-efficiency agent-evaluationSkill GraderSEOAgent

What is this?

The supplied results describe emerging benchmarks and contrastive tests for agent “skills”: instruction packages that coding agents retrieve or load to guide task execution. Reported findings suggest skills can reduce accuracy, slow models, or fail under autonomous selection and distractors, while an automated triage workflow reportedly classifies functional failures and efficiency regressions with 93.6% and 79.7% accuracy respectively. However, the snippets do not establish what the Show HN project “Skill Grader” specifically does, who built it, or whether its own independent testing has produced results; they support the broader hypothesis rather than the particular launch.

Why it matters to Scott

Scott already holds the core position in “Fat AGENTS.md Anti-Pattern” and prescribes contrastive, repeatable gates in “Evaluation-Driven Development.” A credible independent skill grader could operationalize and test that claim across agent skills, but the supplied evidence establishes neither Skill Grader’s implementation nor results, so this is not yet external convergence or a challenge.
ip:concept.fat-agents-md-anti-patternip:concept.evaluation-driven-developmentip:concept.skills-and-workflowsip:framework.context-engineeringradar:concept.agent-skillsradar:concept.agent-evaluationradar:handbook-md-agent-policy-failure
queries asked of Scott's wikis
  • coding-agent skill selection and retrieval failures
  • context bloat versus procedural skill usefulness
  • contrastive evals with-skill versus no-skill
  • automated grading of coding-agent instructions
  • agent harness benchmarks and regression testing
  • skills versus RAG and tool documentation

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 1129h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

08-25 15:15⭐ origin directly observedShow HN: See if bloated 'skills' are degrading your coding agent's performance
aleclindz on hacker news
—
08-26 16:35first on r/ClaudeAI · published · +25.3hCan Claude handle large, long-term projects? After three months, his performance has deteriorated significantly.
Neofaction
—
08-27 18:12first on hacker news · published · +51.0hI measured what my Claude.md, skills and hooks are worth
smartwordworld
—
09-13 13:54first on r/artificial · published · +454.6hCOBRA-Skills: Contextual Bandits for Efficient Agent Skill Optimization (Open Source)
GardenDelicious1476
—
08-25 15:15amplified on hacker newshn.story.49435609
aleclindz
peak 1 · 0 comments · 0% of case engagement
08-26 16:35amplified on r/ClaudeAIreddit.post.1vz26ik
Neofaction
peak 28 · 48 comments · 10% of case engagement
08-26 16:58amplified on r/ClaudeAIreddit.post.1vz2tbv
Clear-Dimension-6890
peak 0 · 7 comments · 1% of case engagement
08-27 01:26amplified on r/ClaudeAIreddit.post.1vzg2hx
mpanase
peak 1 · 2 comments · 0% of case engagement
08-27 18:12amplified on hacker newshn.story.49468945
smartwordworld
peak 8 · 2 comments · 2% of case engagement
08-28 03:34amplified on r/ClaudeAIreddit.post.1w0fhp7
fuckme
peak 1 · 1 comments · 0% of case engagement
36 more amplifiers in ainews.case_chain
08-25 15:21our radar first saw it · +0.1hdiscovery anchor: hn.story.49435609—

Evidence (42) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn ⭐Show HN: See if bloated 'skills' are degrading your coding agent's performancealeclindz10
🟠 redditClaude Code - Taming the Beast
ClaudeAI
Clear-Dimension-689007
🟠 redditCan Claude handle large, long-term projects? After three months, his performance has deteriorated significantly.
ClaudeAI
Neofaction2848
🟠 redditHow to disable corporate "agentic system" plugin
ClaudeAI
mpanase12
🟧 hnI measured what my Claude.md, skills and hooks are worthsmartwordworld82
🟠 redditI ran 6 AI presentation-design skills against one brief and had them review each other: 36 decks, 36 explainers, 9 review passes
ClaudeAI
fuckme11
🟠 redditI built a Claude Code skill that improved analysis coverage in my tests — can someone try to break my results?
ClaudeAI
LibertariansAI10
🟠 redditI started treating CLAUDE.md like code — every rule needs a reason, or it gets deleted
ClaudeAI
halluci_data04
🟠 redditI got tired of AI frontend skills being huge instruction dumps, so I built a registry instead
ClaudeAI
Sea-Firefighter989602
🟠 redditClaude code and breaking up Claude.md in large projects
ClaudeAI
SoCal_Hunter1612
🟠 redditClaude is genuinely good at data engineering now, it just needed the right context loaded in
ClaudeAI
Unknown-33323
🟠 redditClaude is genuinely good at data engineering now, it just needed the right context loaded in
ClaudeAI
Unknown-33311
🟧 hnWhat agent skills are made of: programming languages3Mathematicians20
🟧 hnSkills MCPgengirish44
🟧 hnShow HN: JIT Skill Architecture for AI Agents (Without Context Decay)eshaforostov10
🟧 hnShow HN: Turn repeated coding-agent corrections into rules/skillsblumeCodes20
🟠 redditI built a local verifier for when Claude Code skips required Skills
ClaudeAI
RefrigeratorOwn994121
🟠 redditIs it common for Claude, especially the Opus 5 model, to overdo things when following a detailed prompt?
ClaudeAI
Own-Adhesiveness-705217
🟠 redditAgents Follow Negative Prohibitions, Forget Positive Guidance
ClaudeAI
Wsz2020024
🟠 redditI re-measured my own published skill results with a fixed instrument. None of them separated from noise. Both runs' receipts are public.
ClaudeAI
maverick_man111121
🟠 redditClaude Code ignoring skill files and acting unpredictably
ClaudeAI
CAPSEnthusiast08
🟠 redditWe hit skills sprawl at 33 skills in a 6-person team. Here’s how we fixed it.
ClaudeAI
jpmc_19708
🟠 redditClaude Fable 5.1 shipped. I re-ran my skill evals against it with the skill and suite hashes asserted identical before the first call. Nothing moved beyond noise, and the interesting part is which arm moved.
ClaudeAI
maverick_man111105
🟠 redditWhy static CLAUDE.md files degrade agent performance (and how to automate skill distillation)
ClaudeAI
navune24
🟧 hnClaude Code skills for advanced context engineering techniques and patternsleovs09434
🟠 redditI open-sourced the Claude Code skill I use to edit my own videos
ClaudeAI
ustype012
🟠 redditDo we still need frontend-design skill for Fable 5 / 5.1?
ClaudeAI
Anxious_Sea_18732012
🟧 hnShow HN: Skillsaw – Linter for Contextstbenjam10
🟠 redditAnthropic says "double-check your work" is now an anti-pattern. I counted 125 of those lines in my own config and cannot tell which ones matter.
ClaudeAI
Frequent-Ad-83614462
🟧 hnShow HN: Skillctl – audit context cost and conflicts across your agent skillszongwu23320
🟠 redditWe measured whether our 20 skills actually fire. Baseline recall was 46%, and our first detector only understood Claude's Skill tool.
ClaudeAI
EvalRaccoonDev44
🟠 redditVerbosity in skills
ClaudeAI
mdspan22
🟧 hnAnthropic released a CLI to evaluate skills and pluginsedonadei31
🟠 redditCOBRA-Skills: Contextual Bandits for Efficient Agent Skill Optimization (Open Source)
artificial
GardenDelicious147610
🟠 redditSkill authors: would you put a "works in the field" badge on your SKILL.md? (real-run pass rate, not a linter score)
ClaudeAI
No_Advertising253609
🟧 hnShow HN: Skillzero – save tokens by omitting skills from agent contextkurtextrem22
🟠 redditHow I Actually Use Claude Code
ClaudeAI
lucavallin11
🟧 hnShow HN: Linting 216 public Claude Code skills – 69% won't reliably triggersgharlow30
🟠 redditWhy your skills suck(and how to improve them)
ClaudeAI
Hydronix273101
🟠 redditHas anyone else noticed this? Claude writes prompts in a way that somehow leads to more prompts instead of actually achieving the goal.
ClaudeAI
MoreFaithlessness954013
🟠 redditAgent Dispatcher — Automatically routes tasks to the right role, skills, tools, and context
ClaudeAI
drankthedew2732
🟠 redditThe biggest with-and-without improvement in my skill test came from a skill the Skill tool never called once
ClaudeAI
maverick_man111126

Interpretation history

Decision trace