2026-10-11 16:38 UTC

PromptArmor claims Microsoft Copilot Cowork's AI gateway can be hijacked to bypass sandboxing and exfiltrate local files, and Microsoft's mitigation โ€” or inaction โ€” establishes agent-gateway trust boundaries as a practical attack surface for consumer agent products.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: highagentic-security sandbox-escape ai-infrastructurePromptArmorMicrosoft

What is this?

PromptArmor published a disclosure that Microsoft Copilot Cowork โ€” Microsoft's M365 agent that runs user 'Skills' and code in an internet-locked-down cloud sandbox โ€” could be hijacked via indirect prompt injection to abuse the agent's own AI gateway (the network path the agent itself needs to reach its servers), bypassing the sandbox egress restrictions and giving attackers remote control and file exfiltration. Per PromptArmor's timeline it was reported June 24, 2026 and Microsoft confirmed mitigation on August 19, 2026 โ€” so the vendor response has already landed, roughly six weeks before this case's 'response expected imminently' framing; this is post-fix publication. It sits inside a string of PromptArmor findings against Cowork-class products: file exfiltration via pre-authenticated download links and auto-approved Email/Teams sends, the same attack demonstrated against Anthropic's Claude Cowork by allowlisting the Anthropic API for sandbox egress, and Skills reaching admin-blocked models like DeepSeek through the agent's own access path. The Hacker News thread shows genuine dispute over whether this is a 'sandbox escape' or simply the inherent risk of running untrusted skills (plugins) โ€” one commenter's 'no trust boundary between trusted and untrusted context' vs. 'this is just a malicious extension, caveat emptor' โ€” which is the interesting fault line, alongside Microsoft's own application card claiming Cowork can't touch local files and always requires approval for sensitive actions.

Why it matters to Scott

Dated receipt for the containment thesis his security canon rests on: the sandbox egress boundary failed because the agent's own AI gateway โ€” the one sanctioned network path โ€” became the exfiltration channel, exactly his confused-deputy/tool-gateway claim in Sandboxed Execution and SiloOS, and the HN dispute ('sandbox escape' vs 'just a malicious extension') is his containment-vs-provenance split playing out in public. Design transfer is direct: his own stacks carry the same trusted model-gateway path (OpenClaw's ve1 gateway, ask's LAN LiteLLM proxy, AWS-Firewall-scoped egress), and Microsoft's falsified app-card assurances are fresh receipts for Architecture, Not Vibes and the provenance stack's signed-skills/authority argument.
ip:concept.sandboxed-executionip:source.siloosip:framework.agent-provenance-stackip:concept.confused-deputy-problemip:framework.architecture-not-vibesdev:project.silo-osdev:technology.aws-network-firewallradar:concept.agent-sandboxingradar:concept.prompt-injectionradar:concept.data-exfiltrationradar:openai-dns-sandbox-escaperadar:claude-cowork-sharedroot-sandbox-escaperadar:atlassian-rovo-prompt-injection-exfiltrationradar:anthropic-skill-scanner-backdoor-bypass
queries asked of Scott's wikis
  • agent harness sandbox design trust boundary between instructions and tool results
  • indirect prompt injection exfiltration channel network egress tool gateway
  • untrusted skills plugins as attack surface โ€” supply chain for agent capabilities
  • human approval gating sensitive actions โ€” automatic vs confirmed agent actions
  • agent memory wiki poisoning โ€” untrusted content becoming trusted context
  • local file access and permission boundaries in coding agents

Measured heat

now 0 pts/hpeak 4 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 263h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-30 19:32 (minted)โญ origin echo-reconstructedHijacking Copilot Cowork's AI Gateway to Bypass Sandboxing and Exfiltrate Files.
PromptArmor on blog (echo) ยท attributed from hn.story.49911869 ยท published time unknown
โ€”
09-30 17:28first on hacker news ยท published ยท lag ?Hijacking Copilot Cowork's AI Gateway to Bypass Sandboxing and Exfiltrate Files
jerryShaker
โ€”
10-07 10:42first on r/artificial ยท published ยท lag ?One Copilot model refused. Another leaked secrets about half the time.
Haunting_Ganache_850
โ€”
09-30 17:28amplified on hacker newshn.story.49911869
jerryShaker
peak 2 ยท 0 comments ยท 23% of case engagement
10-07 10:42amplified on r/artificial ๐Ÿ‘‘reddit.post.1wzt99w
Haunting_Ganache_850
peak 6 ยท 6 comments ยท 77% of case engagement
09-30 18:21our radar first saw it ยท lag ?discovery anchor: hn.story.49911869โ€”
pace: p35 vs 1188 stories at the 168h mark (now 263h old) โ€” ahead of agentgate-signed-agent-receipts (1.3x), behind agentic-determinism-index (0.8x)

Evidence (3) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnHijacking Copilot Cowork's AI Gateway to Bypass Sandboxing and Exfiltrate FilesjerryShaker20
๐ŸŸง echo.blog โญHijacking Copilot Cowork's AI Gateway to Bypass Sandboxing and Exfiltrate Files.PromptArmorโ€”โ€”
๐ŸŸ  redditOne Copilot model refused. Another leaked secrets about half the time.
artificial
Haunting_Ganache_85056

Interpretation history

Decision trace