PromptArmor claims Microsoft Copilot Cowork's AI gateway can be hijacked to bypass sandboxing and exfiltrate local files, and Microsoft's mitigation โ or inaction โ establishes agent-gateway trust boundaries as a practical attack surface for consumer agent products.
state: corroboratedheat: lowuncertainty: mediumconvergesscott: highagentic-security sandbox-escape ai-infrastructurePromptArmorMicrosoft
What is this?
PromptArmor published a disclosure that Microsoft Copilot Cowork โ Microsoft's M365 agent that runs user 'Skills' and code in an internet-locked-down cloud sandbox โ could be hijacked via indirect prompt injection to abuse the agent's own AI gateway (the network path the agent itself needs to reach its servers), bypassing the sandbox egress restrictions and giving attackers remote control and file exfiltration. Per PromptArmor's timeline it was reported June 24, 2026 and Microsoft confirmed mitigation on August 19, 2026 โ so the vendor response has already landed, roughly six weeks before this case's 'response expected imminently' framing; this is post-fix publication. It sits inside a string of PromptArmor findings against Cowork-class products: file exfiltration via pre-authenticated download links and auto-approved Email/Teams sends, the same attack demonstrated against Anthropic's Claude Cowork by allowlisting the Anthropic API for sandbox egress, and Skills reaching admin-blocked models like DeepSeek through the agent's own access path. The Hacker News thread shows genuine dispute over whether this is a 'sandbox escape' or simply the inherent risk of running untrusted skills (plugins) โ one commenter's 'no trust boundary between trusted and untrusted context' vs. 'this is just a malicious extension, caveat emptor' โ which is the interesting fault line, alongside Microsoft's own application card claiming Cowork can't touch local files and always requires approval for sensitive actions.
Why it matters to Scott
Dated receipt for the containment thesis his security canon rests on: the sandbox egress boundary failed because the agent's own AI gateway โ the one sanctioned network path โ became the exfiltration channel, exactly his confused-deputy/tool-gateway claim in Sandboxed Execution and SiloOS, and the HN dispute ('sandbox escape' vs 'just a malicious extension') is his containment-vs-provenance split playing out in public. Design transfer is direct: his own stacks carry the same trusted model-gateway path (OpenClaw's ve1 gateway, ask's LAN LiteLLM proxy, AWS-Firewall-scoped egress), and Microsoft's falsified app-card assurances are fresh receipts for Architecture, Not Vibes and the provenance stack's signed-skills/authority argument.
ip:concept.sandboxed-executionip:source.siloosip:framework.agent-provenance-stackip:concept.confused-deputy-problemip:framework.architecture-not-vibesdev:project.silo-osdev:technology.aws-network-firewallradar:concept.agent-sandboxingradar:concept.prompt-injectionradar:concept.data-exfiltrationradar:openai-dns-sandbox-escaperadar:claude-cowork-sharedroot-sandbox-escaperadar:atlassian-rovo-prompt-injection-exfiltrationradar:anthropic-skill-scanner-backdoor-bypass
queries asked of Scott's wikis
- agent harness sandbox design trust boundary between instructions and tool results
- indirect prompt injection exfiltration channel network egress tool gateway
- untrusted skills plugins as attack surface โ supply chain for agent capabilities
- human approval gating sensitive actions โ automatic vs confirmed agent actions
- agent memory wiki poisoning โ untrusted content becoming trusted context
- local file access and permission boundaries in coding agents
Measured heat
now 0 pts/hpeak 4 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 263h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p35 vs 1188 stories at the 168h mark (now 263h old) โ ahead of agentgate-signed-agent-receipts (1.3x), behind agentic-determinism-index (0.8x)
Evidence (3) โ โญ canonical anchor
Interpretation history
2026-10-07T11:30:16Z
The case shifts from single-researcher disclosure awaiting vendor response to a post-fix publication โ PromptArmor's own timeline shows Microsoft confirmed mitigation Aug 19 โ whose durable content is now class-level: two independent researchers (PromptArmor on Cowork's gateway, Adversa.ai on Copilot CLI) show Copilot agent products exfiltrating secrets through their own sanctioned paths, with safety turning out model-dependent (one model refused, another leaked ~half the time). Attention never ignited anywhere, so heat cools even as the evidential state promotes.
2026-10-07T11:24:17Z
evidence attached: reddit.post.1wzt99w โ Second independent researcher (Adversa.ai) demonstrating a working Copilot agent exfiltration chain with model-dependent resistance โ materially broadens the Copilot attack-surface episode beyond the gateway-hijack claim.
2026-09-30T19:40:00Z
grounded: converges/high โ Dated receipt for the containment thesis his security canon rests on: the sandbox egress boundary failed because the agent's own AI gateway โ the one sanctioned
2026-09-30T19:32:07Z
case created โ Fresh sandbox-escape disclosure by a credible security researcher against a flagship agent product, with a vendor response expected imminently.
Decision trace
- 10-11 16:44review_screenjev screen: no material development (noul=0.22)
- 10-08 05:23sensor_dirtycomment_update
- 10-08 00:59attention_routeThe editor compared this story and chose to keep watching.
- 10-07 22:30repriceThe case shifts from single-researcher disclosure awaiting vendor response to a post-fix publication โ PromptArmor's own timeline shows Microsoft confirmed mitigation Aug 19 โ whose durable conte
- 10-07 22:24attachSecond independent researcher (Adversa.ai) demonstrating a working Copilot agent exfiltration chain with model-dependent resistance โ materially broadens the Copilot attack-surface episode beyond the
- 10-07 22:22propose_attachSecond independent researcher (Adversa.ai) demonstrating a working Copilot agent exfiltration chain with model-dependent resistance โ materially broadens the Copilot attack-surface episode beyond the
- 10-01 05:40groundDated receipt for the containment thesis his security canon rests on: the sandbox egress boundary failed because the agent's own AI gateway โ the one sanctioned network path โ became the exfiltra
- 10-01 05:32createFresh sandbox-escape disclosure by a credible security researcher against a flagship agent product, with a vendor response expected imminently.