2026-10-11 16:37 UTC

The paper’s authors claim attacker-controlled low-privilege material in an AI agent’s context can induce actions using the agent’s higher privileges, requiring provenance-aware context isolation and privilege boundaries beyond prompt or tool-call filtering.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: highagentic-security context-security prompt-injection

What is this?

The case's namesake paper is now traceable as arXiv 2609.01222, "What's in Your Agent's Context?": it taxonomizes context privilege escalation into Message-Role CPE (attacker-controlled low-privilege content incorporated into a higher-privileged message role) and Cross-Scope CPE (attacker content persisting beyond its original context), and claims a systemic analysis of 12 real-world harnesses including Claude Code and Codex. Authorship is still not established in the supplied snippets, and the empirical findings are known only from the paper's own abstract, so the experiments and results remain unverified. Independent industry work now corroborates the broader mechanism: Microsoft disclosed prompt-injection-to-RCE in agent frameworks ('when prompts become shells', CVE-2026-26030), Cyera published four OpenClaw CVEs framing the agent as 'the attacker's execution layer', a survey citing Trail of Bits reports prompt-injection-to-RCE escalation bypassing only 1–2 approval layers on code agents, and coverage of a lab-hosted 'Agents of Chaos' effort (agentsofchaos.baulab.info — possibly but not confirmed the same paper) cites OpenClaw CVE-2026-27001, where an unsanitized working directory embedded in the system prompt became an injection channel.

Why it matters to Scott

Consequential other parties have now independently arrived where Scott argued first: Microsoft's prompt-injection-to-RCE disclosure, Cyera's OpenClaw CVEs framing the agent as 'the attacker's execution layer', and a Trail of Bits-cited report that code-agent escalation clears only 1–2 approval layers all corroborate the paper's core claim that filtering-class defenses are insufficient and provenance/privilege boundaries are required — dated-receipts material for the Agent Provenance Stack and SiloOS even though the paper's own experiments remain unverified. The newly traceable taxonomy also maps directly onto his canon: Message-Role CPE formalises the role-label-as-authority failure of his Chat Era Trust Model and confused-deputy work, and Cross-Scope CPE extends it — this strengthens and receipts his thesis rather than merely illustrating it.
ip:framework.agent-provenance-stackip:concept.confused-deputy-problemip:concept.chat-era-trust-modelip:concept.taint-trackingip:framework.siloosip:concept.guardrail-illusionradar:concept.prompt-injectionradar:concept.privilege-escalationradar:concept.agent-provenanceradar:concept.agent-authorizationradar:claude-code-ghost-user-messagesradar:atlassian-rovo-prompt-injection-exfiltrationradar:mcp-tool-sequence-guardrail-bypassradar:repository-content-agent-injection
queries asked of Scott's wikis
  • confused deputy prompt injection authority
  • context provenance trust levels message roles
  • Agent Provenance Stack taint tracking design
  • SiloOS agent privilege isolation
  • prompt injection detector filtering limits
  • least privilege agent tool approval boundaries

Measured heat

now 0 pts/hpeak 5 pts/hcomments 0/hpeers p16momentum: steady3 platformsage 945h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-02 07:26 (minted)⭐ origin echo-reconstructedThe paper identifies context privilege-escalation attacks against AI agents as a distinct security threat arising from what enters an agent’
paper authors on paper (echo) · attributed from hn.story.49532763 · published time unknown
—
09-02 07:07first on hacker news · published · lag ?What's in Your Agent's Context? Context Privilege Escalation Attacks Against AI
sbulaev
—
09-02 11:35first on r/ClaudeAI · published · lag ?Well I almost got prompt injected
autistamine
—
09-12 21:26first on r/artificial · published · lag ?Spike traps for AI: fun food for thought
Street_Estate2342
—
09-17 01:40first on r/OpenAI · published · lag ?From OpenAI: An unreleased research model inserted unrelated instructions, including instructions to disregard its normal constraints, into summaries used to continue its work in a new context window.
Rare_Guide_9830
—
09-20 20:12first on r/LocalLLaMA · published · lag ?gemma 4 ran a shell command a web page told it to. qwen 3.8 mostly didn't.
divinetribe1
—
09-02 07:07amplified on hacker newshn.story.49532763
sbulaev
peak 1 · 0 comments · 0% of case engagement
09-02 11:35amplified on r/ClaudeAI 👑reddit.post.1w57t43
autistamine
peak 884 · 68 comments · 59% of case engagement
09-02 16:56amplified on hacker newshn.story.49539111
Miyamura80
peak 2 · 0 comments · 0% of case engagement
09-04 19:00amplified on hacker newshn.story.49568736
ta988
peak 1 · 0 comments · 0% of case engagement
09-06 05:53amplified on hacker newshn.story.49583720
aghuang
peak 1 · 0 comments · 0% of case engagement
09-07 11:44amplified on hacker newshn.story.49597166
edf13
peak 3 · 0 comments · 0% of case engagement
19 more amplifiers in ainews.case_chain
09-02 07:21our radar first saw it · lag ?discovery anchor: hn.story.49532763—
pace: p93 vs 519 stories at the 720h mark (now 945h old) — ahead of claude-mods-in-process-extensions (1.0x), behind swe-bench-pro-harness-cost-parity (1.0x)

Evidence (26) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnWhat's in Your Agent's Context? Context Privilege Escalation Attacks Against AIsbulaev10
🟧 echo.paper ⭐The paper identifies context privilege-escalation attacks against AI agents as a distinct security threat arising from what enters an agent’paper authors——
🟠 redditWell I almost got prompt injected
ClaudeAI
autistamine88268
🟧 hnShow HN: Sleeper Agents in Robot Dogs and Kinetic Prompt InjectionsMiyamura8020
🟧 hnNetworkManager Works to Enforce AI Policy by Tricking AI Agents to Add a Canaryta98810
🟧 hnRevealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMsaghuang10
🟧 hnAn agent skill can hand a stranger your shell – hours after you installed itedf1330
🟧 hnAgent pasted a fake captcha into Terminal, it lived on my Mac for 30 daysgatreddi10
🟧 hnTaint tracking on AI agent traces: 0.48 precision on AgentDojoCodeNinja77830
🟠 redditWe told our own agent to betray us: 24-tool email MCP server, and what the approval gate had to survive
ClaudeAI
Soft-Lie-43405
🟧 hnA Git Config Key Ran Code in Seven Coding Agents. The 2022 Fix Does Not Stop Itedf1340
🟠 redditWhy I am receiving a prompt injection with <system-reminder> on Claude Code ?
ClaudeAI
Greenerli39
🟠 redditSpike traps for AI: fun food for thought
artificial
Street_Estate234244
🟧 hnA computational constitution to stop LLM agents from bricking serversmisqe32
🟧 hnPyshackle: A hard pre-execution gate for AI agent tool calls (open source)SHACKLE-PRO-10
🟧 hnWe gave a coding agent the lethal trifecta: data, internet, a public repojoeyorlando111
🟧 hnShow HN: What an agent does when anyone can read and rewrite its contextljedrz30
🟧 hnShow HN: Open-Source Lightweight Prompt Injection SafetyHiskias10
🟧 hnGitLost: We Tricked GitHub's AI Agent into Leaking Private Reposfagnerbrack11
🟧 hnI caught an LLM-powered recruiter with a prompt injection on LinkedInbucket201511
🟠 redditFrom OpenAI: An unreleased research model inserted unrelated instructions, including instructions to disregard its normal constraints, into summaries used to continue its work in a new context window.
OpenAI
Rare_Guide_983015277
🟧 hnOpenAI models secretly generate instructions to ignore constraintstheahura12537
🟠 redditgemma 4 ran a shell command a web page told it to. qwen 3.8 mostly didn't.
LocalLLaMA
divinetribe1612
🟧 hnCan open-source prompt-injection detectors catch realistic AI agent attacks?northbridgedev84
🟧 hnAI agent memory can be poisoned – and later treated as the user's own pastrodicarsone20
🟧 hnI tested whether AI agents obey text you never see. 4 of 8 did, every timenickcosta10

Interpretation history

Decision trace