2026-10-11 17:11 UTC

PromptArmor claims crafted backdoored agent skills can evade Anthropic’s skill scanner while retaining malicious behavior, exposing a supply-chain gap that would require stronger artifact verification or runtime isolation.

state: expiredheat: lowuncertainty: highconvergesscott: highagentic-security agent-harnesses coding-agentsPromptArmorAnthropic

What is this?

PromptArmor reports that AI-agent “skills”—instruction artifacts that can access packages, URLs, and external services—create a software supply-chain risk through prompt injection, remote instruction loading, data exposure, and malicious execution. The case specifically claims a crafted backdoored skill passed Anthropic’s skill scanner while preserving attacker-controlled runtime behavior, implying that static or LLM-based review may not establish an artifact’s safety. The supplied snippets support the broader scanner-evasion and dependency-risk pattern, but provide little direct detail or independent corroboration of PromptArmor’s Anthropic-specific experiment.

Why it matters to Scott

PromptArmor’s claimed scanner bypass directly supports Scott’s load-bearing position that model/static review cannot establish runtime safety: agent extensions require hard containment, least-authority execution, and verifiable provenance. If independently reproduced, the Anthropic-specific result is a strong dated-receipts and publishing opportunity for Architecture, Not Vibes, the Agent Provenance Stack, and SiloOS; current evidence remains largely PromptArmor testimony.
ip:framework.architecture-not-vibesip:framework.agent-provenance-stackdev:project.silo-osdev:concept.padded-cell-agent-architectureradar:skillpreflight-agent-skill-scoringradar:augur-hidden-content-scannerradar:concept.agent-skillsradar:concept.software-supply-chain
queries asked of Scott's wikis
  • agent skill supply-chain security
  • static scanning vs runtime isolation for agents
  • artifact signing and provenance for agent extensions
  • remote instruction loading in agent harnesses
  • capability sandboxing for coding agents
  • trust boundaries for MCP plugins and skills

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnAnthropic's Skill Scanner Defeated by Backdoored SkillsjerryShaker10
🟧 echo.blog ⭐PromptArmor’s original research says Anthropic’s scanner can pass a backdoored Skill whose runtime-fetched, attacker-controlled data reachesPromptArmor Threat Intelligence Team——

Interpretation history

Decision trace