2026-10-11 17:15 UTC

IronWarden's maintainer claims the released Rust reverse proxy redacts PII from streaming OpenAI/Anthropic/Ollama traffic in microseconds (<0.07 ms p95 routing overhead, self-benchmarked) with sliding-window SSE token rehydration and HMAC-chained audit logs, making inline privacy screening practical without buffering or client changes.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumpii-redaction llm-privacy agentic-security llm-gatewaysSomnerd

What is this?

IronWarden is a newly released open-source Rust 'sovereign AI reverse proxy & privacy firewall' (maintainer: Somnerd) that sits between an application and LLM providers β€” OpenAI, Anthropic, Ollama β€” and redacts PII from traffic inline; the maintainer claims <0.07 ms p95 routing overhead, sliding-window SSE token rehydration so streaming responses are restored without buffering or client changes, and HMAC-chained tamper-evident audit logs, all self-benchmarked and unvalidated. Notably, the supplied web results never surface IronWarden itself: they instead show a crowded, fast-moving niche of functionally near-identical tools β€” Eidolon (Rust; regex+BERT-NER hybrid redaction with AES-256-GCM rehydration mapping and egress DLP), Cloakpipe (Rust; SSE streaming rehydration for RAG, <5 ms), tamga (Go; sub-millisecond, GDPR/KVKK), and several EU-AI-Act gateways with hash-chained audit logs. So the underlying pattern is clearly real and actively contested, but these snippets neither validate IronWarden's specific latency/differentiation claims nor establish anything about the maintainer.

Why it matters to Scott

Independently ships Scott's documented privacy-tokenised agent boundary β€” proxy-side PII redaction before the model, vault-style rehydration at the boundary (the SSE token-rehydration mechanic), and tamper-evident audit β€” making it a dated external receipt for his proxy-mediated tokenisation position. Its sub-0.1 ms self-benchmarked claim bears directly on his latency-accuracy-asymmetry position (whether privacy controls must cost the live clock, or can sit inline on streaming paths) and on infrastructure he actively operates β€” the LiteLLM single-gateway, the Ollama endpoint, and the appliance's fail-closed PII tokenisation. Stays at medium: the numbers are self-reported and unvalidated, the maintainer is unknown, and the radar already tracks four near-identical proxies (LLM-Shield, Cloakwall, Nenya, Preflight) β€” the new content is the extreme-latency claim plus the streaming-rehydration implementation, not the pattern itself.
dev:concept.privacy-tokenized-agent-boundaryip:concept.proxy-mediated-tokenisationip:concept.company-ai-gatewayip:concept.latency-accuracy-asymmetrydev:technology.litellmdev:project.applianceradar:llm-shield-zero-egress-pii-proxyradar:cloakwall-litellm-privacy-auditradar:nenya-provider-secret-redactionradar:preflight-outbound-secret-inspectionradar:concept.llm-gatewaysradar:concept.security-proxiesradar:concept.ai-privacy
queries asked of Scott's wikis
  • LLM proxy PII redaction gateway position
  • SSE streaming token rehydration mapping vault
  • local inference sovereignty vs cloud privacy firewall
  • agent egress DLP secret-scanning harness
  • redaction impact on RAG retrieval quality
  • Rust LLM gateway latency benchmark validation

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 387h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-25 13:26 (minted)⭐ origin echo-reconstructedReleased 'sovereign AI reverse proxy & privacy firewall' in Rust claiming real-time streaming PII redaction, SSE token rehydration, prompt-i
Somnerd on github (echo) Β· attributed from hn.story.49843753 Β· published time unknown
β€”
09-25 12:35first on hacker news Β· published Β· lag ?IronWarden – A microsecond PII firewall for streaming LLMs in safe Rust
nikolasalexandr
β€”
10-05 10:49first on r/ClaudeAI Β· published Β· lag ?I couldn't paste client numbers into Claude, so I built a way for it to make decks without seeing them
shoumikgoswami
β€”
09-25 12:35amplified on hacker newshn.story.49843753
nikolasalexandr
peak 2 Β· 0 comments Β· 18% of case engagement
10-05 10:49amplified on r/ClaudeAI πŸ‘‘reddit.post.1wy5f4d
shoumikgoswami
peak 0 Β· 16 comments Β· 82% of case engagement
09-25 13:21our radar first saw it Β· lag ?discovery anchor: hn.story.49843753β€”
pace: p53 vs 1032 stories at the 336h mark (now 387h old) β€” ahead of astra-skills-prompt-migration (1.1x), behind autobot-persistent-chatgpt-harness (0.9x)

Evidence (3) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnIronWarden – A microsecond PII firewall for streaming LLMs in safe Rust
Retrieved article excerpt

Open article Β· Retrieved 2026-09-25T13:25:20.378371+00:00

# 🏰 IronWarden

### High-Performance Sovereign AI Reverse Proxy & Privacy Firewall

[Rust CI](https://github.com/Somnerd/IronWarden/actions/workflows/rust_ci.yml)
[Release](https://github.com/Somnerd/IronWarden/releases)
[License: MIT](https://github.com/Somnerd/IronWarden/blob/main/LICENSE)
[Docker](https://github.com/Somnerd/IronWarden/pkgs/container/ironwarden)
[Rust](https://github.com/Somnerd/IronWarden/blob/main/Cargo.toml)
[Latency](https://github.com/Somnerd/IronWarden/blob/main/BENCHMARKS.md)

**IronWarden** is a sovereign, ultra-low-latency AI security reverse proxy and PII firewall written in bare-metal Rust.

Point any **OpenAI**, **Anthropic**, or **Ollama/vLLM** SDK client at IronWarden to get **real-time streaming PII redaction**, **sliding-window SSE token rehydration**, **prompt injection defense**, and **cryptographic HMAC-SHA256 audit chaining** β€” with **zero code changes** in your application.

---

## ⚑ Technical Superiority & Latency Benchmark Matrix

| Metric | **IronWarden** (Rust) | **LiteLLM** (Python) | **Portkey** (Node.js) | **Kong AI Gateway** (Lua/Go) |
| --- | --- | --- | --- | --- |
| **Language & Runtime** | Bare-Metal Rust (Tokio/Axum) | Python (FastAPI/Uvicorn) | Node.js (TypeScript) | OpenResty (Lua) / Go |
| **P95 Routing Overhead** | **<0.07 ms** | 18.5 ms | 12.2 ms | 3.4 ms |
| **Streaming PII Redaction** | **Real-Time Sliding Window** | Buffers Entire Stream | Buffers or regex post-hoc | Basic plugin / slow Lua regex |
| **Max Concurrency (1 Core)** | **125,000+ req/s** | ~2,200 req/s | ~4,800 req/s | ~24,000 req/s |
| **Memory Footprint** | **~18 MB** | ~140 MB | ~110 MB | ~85 MB |
| **Data Sovereignty** | **100% Local / On-Prem / VPC** | Local or Cloud | Cloud SaaS Dependent | Self-hosted or Cloud |
| **Audit Log Integrity** | **Cryptographic HMAC-SHA256 Chaining** | Plain Text JSON | Cloud SaaS Dashboard | Standard Access Logs |



---

## πŸ—οΈ Architecture

```
       [ Client / Microservices / OpenAI & Anthropic SDKs ]
                   β”‚
                   β–Ό (HTTP/2, Streaming SSE, JSON-RPC)
       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
       β”‚                   IronWarden Core Gateway                   β”‚
       β”‚                                                             β”‚
       β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
       β”‚  β”‚ Token Bucket     β”‚    β”‚ Axum / Hyper High-Concurrency  β”‚ β”‚
       β”‚  β”‚ GCRA Rate Limit  │───▢│ Non-Blocking Connection Pool   β”‚ β”‚
       β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
       β”‚                                     β”‚                       β”‚
       β”‚                                     β–Ό                       β”‚
       β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
       β”‚  β”‚ Streaming SSE Rehydration Engine                       β”‚ β”‚
       β”‚  β”‚  β€’ Sliding-window token reassembly across chunk splits β”‚ β”‚
       β”‚  β”‚  β€’ Zero-copy string normalization & homoglyph defense  β”‚ β”‚
       β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
       β”‚                                     β”‚                       β”‚
       β”‚                                     β–Ό                       β”‚
       β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
       β”‚  β”‚ Multi-Tier PII & Security Gating                       β”‚ β”‚
       β”‚  β”‚  β€’ Layer 1: SIMD-Accelerated Aho-Corasick Regex Rules  β”‚ β”‚
       β”‚  β”‚  β€’ Layer 2: ShadowNer Named Entity Recognition         β”‚ β”‚
       β”‚  β”‚  β€’ Layer 3: Prompt Injection & Smuggling Guardrail     β”‚ β”‚
       β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
       β”‚                                     β”‚                       β”‚
       β”‚                                     β–Ό                       β”‚
       β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
       β”‚  β”‚ Tamper-Proof Audit Chaining (HMAC-SHA256 Merkle Chain) β”‚ β”‚
       β”‚  β”‚  β€’ Verifiable cryptographic audit trail for EU AI Act  β”‚ β”‚
       β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                       β”‚ (Redacted Outbound TX)
                                       β–Ό
                 [ Upstream LLMs: OpenAI / Anthropic / Local Ollama ]
```

---

## ⚑ Zero-Friction Quickstart

### 1. Run with Docker (1-Command Instant Start)

Spin up IronWarden in 5 seconds with zero configuration:

```
docker run -d --name ironwarden \
  -p 8080:8080 \
  -e UPSTREAM_LLM="https://api.openai.com" \
  -e WARDEN_MODE="hybrid" \
  ghcr.io/somnerd/ironwarden:latest
```

### 2. Verify with Streaming Curl

Send an LLM prompt containing sensitive PII and observe instant streaming token restoration with zero telemetry leakage:

```
curl -X POST http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "gpt-4o",
    "messages": [
      {"role": "user", "content": "Process payment for John Doe, SSN 000-12-3456, IBAN GR1201101250000000012345678."}
    ],
    "stream": true
  }'
```

### 3. Deploy with Docker Compose

```
docker compose up -d
```

### 4. Deploy to Kubernetes with Helm

```
helm install ironwarden ./deploy/helm/ironwarden \
  --set secrets.wardenPepper="0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef" \
  --set secrets.openaiApiKey="sk-..."
```

### 5. Build & Run from Source

```
git clone https://github.com/Somnerd/IronWarden.git
cd IronWarden

# Optional: Download ONNX NER weights (falls back to high-speed heuristic mode if omitted)
./scripts/setup_models.sh

# Run the gateway
export WARDEN_PEPPER=$(openssl rand -hex 16)
cargo run --release -p app
```

---

## πŸ”Œ Universal Drop-in SDK Compatibility

### OpenAI Python SDK

Simply set `base_url` to IronWarden's gateway endpoint:

```
from openai import OpenAI

# Point client to IronWarden β€” zero code modifications required
client = OpenAI(
    api_key="sk-mock-or-real",
    base_url="http://localhost:14141/v1",
    default_headers={"Authorization": "Bearer <your-jwt-or-key>"}
)

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[
        {"role": "user", "content": "Patient John Doe (SSN: 123-45-6789) shows elevated blood pressure."}
    ],
    stream=True # Streaming supported natively with real-time SSE token rehydration
)

for chunk in response:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

# βœ… PII scrubbed before reaching upstream LLM
# βœ… HMAC-chained tamper-evident audit record logged
# βœ… PII seamlessly restored in the output stream
```

### Anthropic Claude Python SDK

```
import anthropic

client = anthropic.Anthropic(
    api_key="sk-ant-...",
    base_url="http://localhost:14141",
    default_headers={
        "Authorization": "Bearer <your-jwt-or-key>",
        "X-IronWarden-Upstream-Key": "sk-ant-..."
    }
)

message = client.messages.create(
    model="claude-3-5-sonnet-20241022",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Customer Jane Smith (Email: [email protected]) requested a refund."}]
)
print(message.content[0].text)
```

### Dynamic Upstream Routing Headers

| Header | Description | Default |
| --- | --- | --- |
| `X-IronWarden-Target-URL` | Explicitly overrides upstream URL per-request (e.g. `http://localhost:11434/v1/chat/completions`) | Inferred from model name |
| `X-IronWarden-Upstream-Key` | Per-request API key for upstream provider | `OPENAI_API_KEY` / `ANTHROPIC_API_KEY` |

**Automatic Model Routing:**

- `claude-*` βž” Anthropic API (`https://api.anthropic.com/v1/messages`)
- `llama*`, `mistral*`, `phi*`, `gemma*`, `qwen*` βž” Local Ollama (`http://localhost:11434/v1/chat/completions`)
- All other models βž” OpenAI API (`https://api.openai.com/v1/chat/completions`)

---

## πŸ›‘οΈ Core Capabilities & Invariants

### 1. Real-Time Streaming SSE Token Rehydration

Unlike standard proxies that buffer the entire response to replace tokens (introducing massive latency and breaking streaming UI), IronWarden implements an **asynchronous SSE sliding-window state machine** (`SseRehydrator`). It dynamically stitches split tokens across partial HTTP chunks in under **0.04 ms** per chunk.

### 2. Hybrid Intelligence PII Shield

- **Deterministic Layer (Aho-Corasick + Entropy Smuggling Protection)**: Ultra-fast regex and entropy heuristics for Credit Cards, SSNs, Emails, Phone Numbers, IBANs, and International IDs (including Greek AMKA/AFM and EU identifiers).
- **Probabilistic Layer (Local ONNX NER)**: In-process DistilBERT Named Entity Recognition for contextual Names, Organizations, and Locations.

### 3. Cryptographic Audit Vault & Strict Fail-Closed Invariants

- **AES-256-GCM Encryption**: Prompt and redaction records are encrypted at rest using your cryptographic pepper.
- **HMAC-SHA256 Hash Chaining**: Every log entry is cryptographically linked to the previous record with continuous full-chain integrity walk verification.
- **Fail-Closed Security**: If storage fills up or audit logging fails, IronWarden physically halts upstream egress to prevent un-audited data leakage.

### 4. Turnkey Compliance Presets

Pre-configured, zero-touch regulatory rule sets ready to deploy:

- **Middle East & GCC Sovereignty** (`config/rules/me.yaml`): Saudi Arabia PDPL (SDAIA), UAE Federal Decree-Law No. 45/2021, Qatar. Emirates ID, Saudi National ID/Iqama, Saudi & UAE IBANs, GCC mobile numbers, Arabic name heuristics.
- **East Asia Sovereignty** (`config/rules/east_asia.yaml`): China PIPL / CSL, Japan APPI, South Korea PIPA, Singapore PDPA. China Resident ID, USCC, China Mobile, Japan My Number, Korea RRN, Singapore NRIC.
- **India DPDP Act 2023** (`config/rules/in.yaml`): PAN cards, Aadhaar numbers, GSTIN, Voter ID (EPIC), Indian Passports, Indian Mobile.
- **GDPR & European Sovereignty** (`config/rules/eu.yaml`, `config/rules/gr.yaml`): EU & Greek national IDs (AMKA, AFM), EU IBANs, Passports, Driving Licenses.
- **HIPAA** (`config/rules/rules_medical.yaml`): Medical records, Patient IDs, MRNs, SSNs.
- **PCI-DSS** (`config/rules/rules.yaml`): Primary Account Numbers (PANs), CVVs, track data.

### 5. Model Context Protocol (MCP) Server

IronWarden includes a native JSON-RPC 2.0 stdio MCP server for agentic AI architectures (Claude Desktop, Cursor, AI agents) with session isolation and prompt sanitization tools:

- `mcp_sanitize_prompt`
- `mcp_restore_prompt`
- `mcp_get_compliance_report`

---

## πŸ“Š Performance Benchmarks

Measured using [Criterion.rs](https://github.com/bheisler/criterion.rs) with 1,000+ iterations per sample. See [BENCHMARKS.md](https://github.com/Somnerd/IronWarden/blob/main/BENCHMARKS.md) for full methodology.

| Metric | Measured Value | Real-World Impact |
| --- | --- | --- |
| **Ingress PII Scrubbing + Shield** | **0.38 ms** (p50) / **1.12 ms** (p95) | <0.1% of standard LLM TTFT |
| **Streaming SSE Rehydration (per chunk)** | **0.04 ms** (p50) / **0.12 ms** (p95) | Zero perceived token streaming stutter |
| **AES-256-GCM + HMAC Audit Persistence** | **0.15 ms** (p50) / **0.42 ms** (p95) | Fully offloaded & asynchronous |
| **Total Added Gateway Overhead** | **< 1.8 ms** (p95) | **< 1.2% total added latency** |
| **Throughput (Single Process)** | **14,200+ req/s** | Scales linearly with CPU cores |
| **Base Memory Footprint** | **~28.4 MB RSS** | Ultra-lightweight edge deployment |



---

## πŸ“ˆ Observability & Grafana Dashboard

IronWarden includes native, production-grade observability:

- **Prometheus Metrics**: `GET /metrics` exposes request counts, blocked prompt injections, redacted PII entities, and available concurrency permits.
- **Turnkey Grafana Dashboard**: `GET /grafana/dashboard` exports the pre-configured Grafana dashboard JSON.
- **Structured Health Inspection**: `GET /health` returns JSON uptime, permit availability, and system status.

### 1-Command Monitoring Stack

Launch IronWarden + Promethe
nikolasalexandr20
🟧 echo.github ⭐Released 'sovereign AI reverse proxy & privacy firewall' in Rust claiming real-time streaming PII redaction, SSE token rehydration, prompt-iSomnerdβ€”β€”
🟠 redditI couldn't paste client numbers into Claude, so I built a way for it to make decks without seeing them
ClaudeAI
shoumikgoswami016

Interpretation history

Decision trace