PromptArmor claims crafted backdoored agent skills can evade Anthropic’s skill scanner while retaining malicious behavior, exposing a supply-chain gap that would require stronger artifact verification or runtime isolation.
state: expiredheat: lowuncertainty: highconvergesscott: highagentic-security agent-harnesses coding-agentsPromptArmorAnthropic
What is this?
PromptArmor reports that AI-agent “skills”—instruction artifacts that can access packages, URLs, and external services—create a software supply-chain risk through prompt injection, remote instruction loading, data exposure, and malicious execution. The case specifically claims a crafted backdoored skill passed Anthropic’s skill scanner while preserving attacker-controlled runtime behavior, implying that static or LLM-based review may not establish an artifact’s safety. The supplied snippets support the broader scanner-evasion and dependency-risk pattern, but provide little direct detail or independent corroboration of PromptArmor’s Anthropic-specific experiment.
Why it matters to Scott
PromptArmor’s claimed scanner bypass directly supports Scott’s load-bearing position that model/static review cannot establish runtime safety: agent extensions require hard containment, least-authority execution, and verifiable provenance. If independently reproduced, the Anthropic-specific result is a strong dated-receipts and publishing opportunity for Architecture, Not Vibes, the Agent Provenance Stack, and SiloOS; current evidence remains largely PromptArmor testimony.
ip:framework.architecture-not-vibesip:framework.agent-provenance-stackdev:project.silo-osdev:concept.padded-cell-agent-architectureradar:skillpreflight-agent-skill-scoringradar:augur-hidden-content-scannerradar:concept.agent-skillsradar:concept.software-supply-chain
queries asked of Scott's wikis
- agent skill supply-chain security
- static scanning vs runtime isolation for agents
- artifact signing and provenance for agent extensions
- remote instruction loading in agent harnesses
- capability sandboxing for coding agents
- trust boundaries for MCP plugins and skills
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-09-04T18:24:06Z
The Anthropic-specific bypass remains an uncorroborated PromptArmor claim, with no reproduction, technical artifact, or vendor response after the validation window. The episode has faded and should reopen only on substantive new evidence.
2026-09-02T17:58:14Z
The validation gap persists: no independent reproduction, exploit artifact, or Anthropic response has appeared. The claim remains relevant but single-source and has not developed enough to merit further attention until substantive evidence arrives.
2026-08-31T17:35:16Z
No independent reproduction, technical artifact, or Anthropic response has emerged; the case remains a consequential but single-source exploit claim. The unchanged observation adds no new meaning, so attention should cool pending validation.
2026-08-31T17:32:50Z
grounded: converges/high — PromptArmor’s claimed scanner bypass directly supports Scott’s load-bearing position that model/static review cannot establish runtime safety: agent extensions
2026-08-31T17:29:09Z
origin walked (codex/luna, conf 0.98): anchor hn.story.49512263 -> echo.blog.5e20c1e077 by PromptArmor Threat Intelligence Team
2026-08-31T17:27:55Z
case created — The report presents a concrete bypass of a named agent-skill security control with direct implications for skill supply-chain defenses.
Decision trace
- 09-05 04:24expireThe Anthropic-specific bypass remains an uncorroborated PromptArmor claim, with no reproduction, technical artifact, or vendor response after the validation window. The episode has faded and should re
- 09-05 04:24alert_silentThe only delta is elapsed time without corroboration; no new event or evidence warrants Scott's attention.
- 09-05 04:24alert_routeThe only delta is elapsed time without corroboration; no new event or evidence warrants Scott's attention.
- 09-03 03:58repriceThe validation gap persists: no independent reproduction, exploit artifact, or Anthropic response has appeared. The claim remains relevant but single-source and has not developed enough to merit furth
- 09-03 03:58alert_silentThe only trigger is staleness, with no consequential new evidence or event. Revisit if Anthropic responds, technical materials are released, or an independent party reproduces the bypass.
- 09-03 03:58alert_routeThe only trigger is staleness, with no consequential new evidence or event. Revisit if Anthropic responds, technical materials are released, or an independent party reproduces the bypass.
- 09-01 03:35repriceNo independent reproduction, technical artifact, or Anthropic response has emerged; the case remains a consequential but single-source exploit claim. The unchanged observation adds no new meaning, so
- 09-01 03:35alert_silentThere is no consequential new delta beyond the already assessed PromptArmor testimony; unchanged engagement does not justify another alert. Revisit if Anthropic responds, exploit materials appear, or
- 09-01 03:35alert_routeThere is no consequential new delta beyond the already assessed PromptArmor testimony; unchanged engagement does not justify another alert. Revisit if Anthropic responds, exploit materials appear, or
- 09-01 03:33alert_shadowPromptArmor has published a concrete exploit claim: a Skill passed Anthropic’s scanning, fetched attacker-controlled runtime data into an unsanitized shell command, and reportedly exfiltrated user fil
- 09-01 03:33alert_routePromptArmor has published a concrete exploit claim: a Skill passed Anthropic’s scanning, fetched attacker-controlled runtime data into an unsanitized shell command, and reportedly exfiltrated user fil
- 09-01 03:32groundPromptArmor’s claimed scanner bypass directly supports Scott’s load-bearing position that model/static review cannot establish runtime safety: agent extensions require hard containment, least-authorit
- 09-01 03:29promote_anchororigin walk conf 0.98
- 09-01 03:27createThe report presents a concrete bypass of a named agent-skill security control with direct implications for skill supply-chain defenses.