2026-10-11 17:13 UTC

OpenAPPA's authors claim their released MIT-licensed deterministic guardrail tracks audience-by-trust data-flow labels outside the agent loop and stops prompt-injection exfiltration with zero successful attacks at 89% task completion on their benchmarks, versus roughly 10% leaks for LLM-judge auto-modes; adoption or independent replication would establish deterministic data-flow guardrails as a practical agent-containment layer.

state: watchingheat: lowuncertainty: mediumconvergesscott: mediumagentic-security agent-containment deterministic-guardrails prompt-injectionMatvey

What is this?

OpenAPPA is an MIT-licensed, open-source agent-security guardrail, currently in preview/RFC stage, that enforces audience-by-trust data-flow labels on everything an agent touches β€” running entirely outside the agent's prompt and execution loop so the model cannot see, negotiate with, or manipulate it. The authors claim it is 100% resistant to prompt-injection (and hallucination-driven) data exfiltration at a cost of only ~1 point of task completion versus an LLM-judge 'auto mode' (89% vs 90% on their benchmarks), with Microsoft's FIDES shown far worse (41% completion / 31% attacks); the case names 'Matvey' as a key person, but the supplied snippets confirm no author identity. All performance numbers are first-party and self-benchmarked β€” no independent replication appears in the supplied material. The surrounding coverage does independently document the gap OpenAPPA targets (Gopher Security's OpenAI Guardrails bypass and CSA's GuardFall report show both LLM-judge guardrails and naive deterministic denylists being defeated), but that same literature is a reminder that deterministic layers have their own bypass history, so the zero-attacks claim should be read as vendor-benchmarked, not established.

Why it matters to Scott

Converges squarely on Scott's canon: OpenAPPA independently ships the exact mechanism his wikis argue for β€” audience-by-trust data-flow labels (taint tracking) enforced by a deterministic membrane running outside the model's loop so the model can't see, negotiate with, or defeat it β€” and its head-to-head receipts (89% completion vs ~10%-leak LLM-judge auto-modes, FIDES comparison) are the first quantified vendor evidence for his architecture-not-vibes / privilege-inversion position, landing in the same territory as his own open-source SiloOS spec. Medium rather than high because everything is first-party and un-replicated β€” it earns watch status under the same independent-testing bar the radar holds Customhouse and Bulwark to, and only becomes a build change (SiloOS comparison/integration) or a LeverageAI governance talking point if adoption or replication lands.
ip:concept.taint-trackingip:framework.siloosip:concept.runtime-containmentip:concept.trust-hierarchyip:concept.privilege-inversionip:concept.deterministic-ai-pendulumdev:project.silo-osdev:concept.padded-cell-agent-architectureradar:customhouse-mcp-exfiltration-proxyradar:bulwark-agent-security-gatewayradar:agent-context-privilege-escalationradar:llm-judge-omission-blindnessradar:concept.agent-containmentradar:concept.agent-isolationradar:concept.prompt-injection-defense
queries asked of Scott's wikis
  • agent harness out-of-band control plane hooks
  • coding agent tool permissioning egress exfiltration controls
  • deterministic validation vs LLM-judge tradeoffs
  • information-flow taint labels agent containment
  • agent-maintained wiki write trust boundaries
  • local agent runtime sandboxing security posture

Measured heat

now 0 pts/hpeak 19 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 1850h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

07-26 14:00⭐ origin echo-reconstructedarXiv:2607.24625 [cs.CR], "APPA: Recoverable Information-Flow Control for Real-World LLM Agents", submitted 27 Jul 2026 (v1), revised 26 Aug
Arseny Kravchenko, Vadim Liventsev, Innokentii Konstantinov, Ildar Iskhakov, Matvey Kukuy (the APPA/OpenAPPA team, Archestra.ai) on paper (echo) Β· attributed from hn.story.49877515
β€”
09-28 13:20first on hacker news Β· published Β· +1535.3hShow HN: OpenAPPA – open-source deterministic guardrails that don't break agents
motakuk
β€”
09-28 13:20amplified on hacker news πŸ‘‘hn.story.49877515
motakuk
peak 25 Β· 12 comments Β· 73% of case engagement
09-30 19:57amplified on hacker newshn.story.49913499
simonpure
peak 10 Β· 2 comments Β· 23% of case engagement
10-01 06:23amplified on hacker newshn.story.49918330
wise_blood
peak 2 Β· 0 comments Β· 4% of case engagement
09-28 15:21our radar first saw it Β· +1537.3hdiscovery anchor: hn.story.49877515β€”

Evidence (4) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: OpenAPPA – open-source deterministic guardrails that don't break agents
Retrieved article excerpt

Open article Β· Retrieved 2026-09-28T15:45:17.840409+00:00

# What is OpenAPPA

How to sing β€œOpenAPPA”

OpenAPPA is a frontier deterministic AI guardrail that is 100% resistant to data exfiltration caused by prompt injection or model hallucination, and the first of its kind that doesn't break agents.

It is open, vendor-agnostic, and MIT-licensed.

And yes, it outperforms competitors on benchmarks:

Task completion

OpenAPPA89%

Claude Auto mode90%

FIDES (Microsoft)41%

Attacks that succeeded

OpenAPPA0%

Claude Auto mode10%

FIDES (Microsoft)31%

[Read the full benchmark results β†’](https://www.openappa.com/evaluation)

## Non-deterministic guardrails miss the problem[#](https://www.openappa.com/#non-deterministic-guardrails-miss-the-problem)

The industry's answer to approval fatigue is a second model that judges each tool call: Claude Code's auto mode, Codex's auto-review, and other [auto-modes](https://www.openappa.com/openappa-vs-auto-mode).

By design, they cannot track data flow across tool calls. Because classifiers are prompt-injectable themselves, harnesses hide tool outputs from them, so the judge never sees the data at all.

Because of their probabilistic design, even the best top out at [99.3%](https://openai.github.io/openai-guardrails-python/ref/checks/prompt_injection_detection/): at millions of calls, 0.7% is a lot of breaches.

## Other deterministic guardrails either break agents or don't work[#](https://www.openappa.com/#other-deterministic-guardrails-either-break-agents-or-dont-work)

Blacklisting commands against an LLM is a dead end. Block `rm -rf /` and the model writes a Python one-liner; block `curl` and it pushes to an external git remote. A list of regexes also gives zero visibility into coverage: nobody can verify that it closes every path, or that three mundane tools chained together don't leak.

Rule sets end up either so tight they break the agent or so intricate nobody can audit what they permit.

## OpenAPPA tracks flows instead of matching patterns[#](https://www.openappa.com/#openappa-tracks-flows-instead-of-matching-patterns)

OpenAPPA is a cross-platform, pluggable engine driven by a [single configuration](https://www.openappa.com/contracts). It runs outside the agent's prompt and execution loop, so the model cannot see, negotiate with, or manipulate it, and it [plugs into an existing agent loop](https://www.openappa.com/add-to-agent) in one place.

appa.tomlversion = 2[[tool]]delta = { … }requires = { … }Coding AgentsLLM ProxiesMCP GatewaysMCP ServersAgents in Production

Instead of allowed and blocked tools, the configuration describes data sources, [audiences](https://www.openappa.com/contracts#audiences), [trust levels](https://www.openappa.com/contracts#trust), and [authorities](https://www.openappa.com/contracts#authorities). Every trajectory carries a security label, `audience Γ— trust`: reading a private repo narrows the audience, reading an unvetted web page lowers trust. The label only ever gets more restrictive, and the engine derives each decision from it algebraically.

An injected prompt telling the agent to leak secrets is irrelevant: you cannot prompt-inject an algebra. Because the configuration is declarative, you can [validate in CI/CD](https://www.openappa.com/validation) that your whole tool graph is covered. The configuration is data-specific, so you can scale to millions of agents without changing it.

## It's full of tricks to help agents accomplish their tasks[#](https://www.openappa.com/#its-full-of-tricks-to-help-agents-accomplish-their-tasks)

Strict enforcement is where utility usually dies: a bare "forbidden" makes an agent stall, retry, and fail. OpenAPPA instead returns a machine-readable [remedy plan](https://www.openappa.com/how-it-works#keeping-agents-useful-under-restrictions) with the ways the agent may legally proceed:

- **Sanitizers** transform the payload, masking secrets or redacting PII, so it can flow to a wider audience. Stock sanitizers ship in the box; custom ones, including model-based ones, plug in with a clear blast radius.
- **Authorities** approve one specific action, through a human or an internal API, without lifting the session's restrictions for later calls.
- **Subagents** isolate an untrusted read in a disposable branch, so the parent trajectory continues unpoisoned.

This is what lifts task completion from 37% to 90% on our [benchmarks](https://www.openappa.com/evaluation) and makes deterministic security practical.
motakuk2512
🟧 echo.paper ⭐arXiv:2607.24625 [cs.CR], "APPA: Recoverable Information-Flow Control for Real-World LLM Agents", submitted 27 Jul 2026 (v1), revised 26 AugArseny Kravchenko, Vadim Liventsev, Innokentii Konstantinov, Ildar Iskhakov, Matvey Kukuy (the APPA/OpenAPPA team, Archestra.ai)β€”β€”
🟧 hnOpenAPPA: Deterministic guardrails that don't break agentssimonpure102
🟧 hnArchestra-AI/OpenAPPA: Deterministic guardrails that don't break agentswise_blood20

Interpretation history

Decision trace