2026-10-11 17:09 UTC

Follow-up research will determine whether self-state attacks can reliably poison persistent agent memory and whether practical integrity defenses can prevent harmful downstream behavior.

state: expiredheat: lowuncertainty: highconvergesscott: highagent-memory prompt-injection agent-securitySchmidhuber

What is this?

“Self-state attacks” are described as attacks that corrupt an AI agent’s persistent memory, instructions, or configuration through otherwise legitimate interactions, potentially influencing later behavior. The supplied research snippets establish active work on memory-poisoning attacks and defenses for memory-based LLM agents, including reliability-conditioned updates, provenance caps, and protections against poisoned retrieved content. However, the snippets do not establish attack reliability, practical defense effectiveness, or Schmidhuber’s role in this work; those points require follow-up research.

Why it matters to Scott

The research independently formalises a failure mode Scott already treats as load-bearing: persistent agent state is an untrusted mutation surface requiring provenance, taint separation, gated writes, audit trails, and rollback. Its eventual attack and defense measurements could directly validate or challenge the security architecture of his agent-maintained wiki and persistent-memory systems, creating a strong dated-receipts opportunity rather than merely another generic prompt-injection example.
ip:concept.durable-external-stateip:framework.agent-provenance-stackip:concept.taint-trackingip:concept.fabricated-history-attackdev:concept.deterministic-agent-control-planedev:concept.propose-finalise-gatedev:project.dev-wikiradar:concept.agent-memoryradar:agenthelm-versioned-agent-memoryradar:karpathy-llm-wiki-adoptionradar:vercel-deepsec-agent-security
queries asked of Scott's wikis
  • persistent agent memory integrity and trust boundaries
  • prompt injection through memory writes
  • provenance-aware memory updates and rollback
  • agent-maintained wiki poisoning defenses
  • separating instructions observations and persistent state
  • memory write permissions validation and audit trails

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (16) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn'Self-State Attacks' New Threat: AI Agents Poisoned via Their Memoryadu_onemore20
🟧 echo.paper ⭐The paper introduces “self-state attacks,” where an agent’s persistent memory, instruction, or configuration is corrupted through legitimateYimeng Chen, Nathanaël Denis, Roberto Di Pietro, Jürgen Schmidhuber——
🟧 hnShow HN: Veracium – agent memory keeping third-party claims from becoming factsqspencer20
🟧 hnShow HN: ZizkaDB: An audit log and state engine database built for AI agentsArshad-Talpur10
🟧 hnShow HN: Open-source, Long-horizon cite-able memory for multi-agent systemsavijeetsingh1630
🟧 hnI built agent memory that retains poisoned data instead of deleting itwalldad220
🟧 hnLLM memory doesn't only get written wrong, it goes wrong latermnzrdev10
🟠 redditClaude has wrong memory of me
ClaudeAI
AggressiveAd52241012
🟧 hnHOM-AIMOS – cryptographically auditable persistent memory for agentswalldad210
🟠 redditIs there a way to stop OpenAI from adding random stuff I asked about to my memory without losing my entire memory?
OpenAI
commandrix147
🟠 redditYour Claude memory files are lying to you and you don't know it
ClaudeAI
Glittering-Agency986010
🟧 hnGit blame for your agent's memorypaulkyle10
🟠 redditYour agent's memory is a file anyone can quietly edit. We published an open protocol that makes that detectable.
OpenAI
zgivod07
🟠 redditA week ago I gave Claude a domain and let it build whatever it wanted. It built a society.
ClaudeAI
zgivod048
🟧 hnWe put llms.txt on 83 websites. In 12 weeks OpenAI's crawler read it 7 timesjamesmackie10
🟠 redditWeird error message
ClaudeAI
bjtak12

Interpretation history

Decision trace