2026-10-11 17:12 UTC

Joshua Penman claims Semantic Overlays can steer a frozen model through trained adapters and raise Qwen 3.5 9B to state-of-the-art results on tested black-box prompt-injection benchmarks, offering a model-level alternative to text-centric guardrails.

state: expiredheat: lowuncertainty: highconvergesscott: mediumprompt-injection agentic-security model-guardrailsJoshua Penman

What is this?

Joshua Penman presents Semantic Overlays, a prompt-injection mitigation technique that applies small learned adapters at selected prefill positions in a frozen model’s residual stream. The overlays create an out-of-band annotation channel for span identity that tokens cannot reproduce, and are described as trained, adaptable, and selectively applied; an arXiv entry lists a paper, interactive demo, code, and released adapters. The supplied snippets do not independently substantiate the specific Qwen 3.5 9B result, the claimed state-of-the-art benchmark performance, or the proposed “NX bit” analogy, so those remain author claims here.

Why it matters to Scott

Semantic Overlays converges with Scott’s taint-tracking and Chat Era Trust Model positions by moving source/span identity into an out-of-band channel that injected tokens cannot imitate. If independently validated it could become a useful preventive layer in his agent-security stack, but benchmark gains would not overturn his Guardrail Illusion claim: adapter steering remains probabilistic model behaviour rather than a deterministic permission boundary.
ip:concept.taint-trackingip:concept.chat-era-trust-modelip:concept.guardrail-illusionip:concept.defense-in-depthradar:transformer-spherical-steering-validationradar:concept.prompt-injectionradar:concept.agentic-securityradar:concept.model-architecture
queries asked of Scott's wikis
  • prompt injection beyond text-centric guardrails
  • trusted data boundaries for agent context
  • model-level steering versus runtime security harnesses
  • adapter-based control of frozen models
  • out-of-band annotations in LLM pipelines
  • defense architecture for tool-using agents

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Semantic Overlays – an NX bit for LLM prompt injection (live demo)joshua_s_penman30
🟧 echo.paper ⭐The original primary artifact is the arXiv paper, submitted August 24, 2026. It introduces “Semantic Overlays”: learned adapters applied at Joshua Penman——

Interpretation history

Decision trace