2026-10-11 16:38 UTC

ghuntley claims the released Preflight proxy inspects complete LLM requests and attachments locally, redacting detected credentials or blocking unsafe requests in enforcing modes before forwarding, adding an inference-boundary exposure control without changing coding harnesses.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: highsecret-redaction llm-gateways agentic-securityghuntley

What is this?

Geoffrey Huntley (ghuntley) has shipped Preflight, a local Rust proxy (127.0.0.1:8081) that sits between any OpenAI-compatible coding harness and the LLM endpoint: it decodes full request bodies and attachments locally (PDF text, OCR, barcodes, metadata), scans them with Gitleaks-derived rules, and by default rewrites detected secrets to stable [REDACTED:rule-id] placeholders before forwarding, with blocking (opaque finding IDs, HTTP 409) and advisory modes also supported. The GitHub docs are unusually thorough for a small release β€” ADRs, structured logging with request correlation, allowlists, and example logs showing sub-30ms policy steps β€” but the supplied snippets show no independent users, audits, or issue traffic, so correctness on real agent traffic and true latency cost remain unproven. The threat model is, however, increasingly corroborated by surrounding coverage: SANS ISC describes hostile or 'free' LLM endpoints harvesting everything coding agents send, a r/LLMDevs thread demonstrates hostile proxies can hijack agents outright, and Cloudanix now sells a commercial on-host 'coding agent guardrail' doing pre-LLM secret/PII redact-or-block β€” alongside fregie's Tokenhush and a third tutorial-style proxy (ogwilliam) as independent implementations. Local pre-inference secret redaction is thus consolidating into a recognized pattern and nascent product category, while Preflight itself remains an unvalidated single-author entry within it.

Why it matters to Scott

ghuntley β€” a consequential agentic-coding figure β€” has independently shipped the inbound airlock of Separation of Powers for Cognition as a runnable local proxy, and this re-ground shows the pattern consolidating into a category (Tokenhush, Cloudanix's commercial on-host guardrail, SANS ISC hostile-endpoint coverage): dated third-party receipts for the ebook, not just one hobby release. It also lands on live infra β€” Preflight occupies exactly the OpenAI-compatible slot in front of a gateway like his LiteLLM, where Presidio already runs, so the concrete actions are (a) evaluating Gitleaks-style secret rules alongside Presidio at that boundary and (b) writing against the design gap between Preflight's irreversible [REDACTED] placeholders and his vault-backed reversible tokenization with hydration at tool boundaries. Caveat carried from the assessment: Preflight itself remains unvalidated (zero adoption, no external scrutiny), so the convergence is pattern-level, not an endorsement of the tool.
ip:framework.separation-of-powers-for-cognitionip:source.separation-of-powers-for-cognition-ebookip:concept.proxy-mediated-tokenisationdev:concept.privacy-tokenized-agent-boundarydev:technology.litellmdev:technology.microsoft-presidioradar:person.ghuntleyradar:underclass-sticky-subscription-poolradar:nenya-provider-secret-redactionradar:cloakwall-litellm-privacy-auditradar:llm-shield-zero-egress-pii-proxyradar:concept.secret-scanningradar:concept.security-proxies
queries asked of Scott's wikis
  • LiteLLM shared gateway secret scanning redaction
  • separation of powers inbound airlock inference boundary
  • coding agent credential leakage into model context
  • reversible tokenization privacy boundary agent requests
  • outbound DLP policy proxy coding harness security
  • ghuntley underclass gateway routing

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p33momentum: steady3 platformsage 505h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-20 15:27 (minted)⭐ origin echo-reconstructedPreflight is a local proxy that scans LLM requests and attachments for secrets, defaults to redaction, supports blocking and advisory modes,
ghuntley on github (echo) Β· attributed from hn.story.49776416 Β· published time unknown
β€”
09-20 14:47first on hacker news Β· published Β· lag ?A proxy that scans LLM requests/attach for secrets before they reach the model
ghuntley
β€”
09-26 17:22first on r/ClaudeAI Β· published Β· lag ?I made a live globe of every server Claude Code talks to, and it blocks the ones you didn't allow
Elegant-Schedule8198
β€”
10-05 14:38first on r/LocalLLaMA Β· published Β· lag ?How do you control what context your coding agent sends to an LLM? I built a local tool to measure and audit it β€” looking for blunt feedback
1982_miguel
β€”
09-20 14:47amplified on hacker newshn.story.49776416
ghuntley
peak 1 Β· 0 comments Β· 9% of case engagement
09-24 03:59amplified on hacker newshn.story.49826094
fregie
peak 3 Β· 0 comments Β· 26% of case engagement
09-26 17:22amplified on r/ClaudeAIreddit.post.1wqw4vc
Elegant-Schedule8198
peak 1 Β· 1 comments Β· 9% of case engagement
10-05 14:38amplified on r/LocalLLaMA πŸ‘‘reddit.post.1wya86v
1982_miguel
peak 0 Β· 7 comments Β· 33% of case engagement
10-09 15:44amplified on r/ClaudeAIreddit.post.1x1oqyl
Otherwise_Ship_9782
peak 0 Β· 5 comments Β· 23% of case engagement
09-20 15:20our radar first saw it Β· lag ?discovery anchor: hn.story.49776416β€”
pace: p36 vs 1032 stories at the 336h mark (now 505h old) β€” ahead of agentgate-signed-agent-receipts (1.3x), behind agent-memory-add-search-evaluation (0.8x)

Evidence (6) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnA proxy that scans LLM requests/attach for secrets before they reach the model
Retrieved article excerpt

Open article Β· Retrieved 2026-09-20T15:22:25.538616+00:00

# preflight

**A local proxy that scans LLM requests and attachments for secrets before they reach the model.**

```
                       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                       β”‚             preflight              β”‚
                       β”‚                                    β”‚
 coding harness ──────►│  /v1/responses                     β”‚
 (any supported        β”‚  /v1/chat/completions              │────► underclass ────► model
  OpenAI-compatible    β”‚  /v1/models                        β”‚
  client)              β”‚                                    β”‚
                       β”‚  decode Β· inspect Β· redact/block   β”‚
                       β”‚  local OCR Β· content-addressed     β”‚
                       β”‚  cache Β· structured logging        β”‚
                       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                             127.0.0.1:8081                     127.0.0.1:8080
```

Someone pastes an `.env` file? Detected credentials become stable placeholders and the request continues. A screenshot or PDF contains a key? preflight decodes it, runs local extraction and OCR, and sanitizes it where supported. When inspection or safe rewriting is impossible, the request stays grounded.

---

## Why

Coding agents read source files, shell output, screenshots, and documents. Secrets can arrive through any of them. preflight sits between the harness and [underclass](https://github.com/ghuntley/underclass), inspecting the outbound copy before inference starts:

- **One base URL change** β€” keep the supported OpenAI HTTP endpoints, tool-call structure, session headers, and response streams.
- **Redact by default** β€” replace detected secrets with `[REDACTED:rule-id]` so ordinary pasted-key incidents do not stop the agent.
- **Attachments get inspected too** β€” decode images and PDFs locally; scan extracted text, metadata, barcodes, and OCR output.
- **No repeated OCR tax** β€” identical attachments reuse completed inspection results while their scope and inspection profile remain the same.
- **No generic randomness detector** β€” prompts and source dumps are entropy soup. Default rules use recognizable credential shapes.
- **Nothing forwarded halfway through inspection** β€” the whole request is checked before it goes to underclass.

## Quick start

With underclass running on `127.0.0.1:8080`, start preflight with its complete document-processing toolchain:

```
nix run github:ghuntley/preflight -- serve
```

Point the harness at:

```
http://127.0.0.1:8081/v1
```

Keep using the existing underclass API key. By default, preflight passes authorization through. You can also configure separate client-facing and upstream credentials.

For a local checkout:

```
devenv shell -- cargo build --locked --bins
devenv shell -- cargo run --locked --bin preflight -- serve
```

Building all binaries also builds the Rust attachment worker. Startup checks the sandboxed native toolchain before the proxy becomes ready.

## How inspection works

- **Parse first.** Walk decoded JSON string leaves, including messages, instructions, tool results, and nested JSON tool arguments. Duplicate object keys are rejected.
- **Match structured content.** Each text unit goes through keyword filtering, a Gitleaks-derived regex, optional entropy filtering, and the trusted allowlist. Ordered text parts and wrapped JWTs get mapped reconstruction passes.
- **Inspect attachments locally.** Resolve their bytes, decode them, extract text, render PDF pages, and run OCR. External attachment URLs never receive the underclass credential.
- **Apply one request-wide decision.** Redact supported spans, rebuild affected artifacts, or reject the request. Rebuilt attachments are inspected again before approval.
- **Forward approved content.** Enforcing modes send the exact approved bytes or sanitized replacement. Unchanged clean requests preserve their original body bytes; underclass's response streams through.

Preflight does not retry inference requests. Underclass owns provider routing and failover.

Design decisions and their trade-offs live in [`docs/adr/`](https://github.com/ghuntley/preflight/blob/main/docs/adr) β€” start with [ADR 0001](https://github.com/ghuntley/preflight/blob/main/docs/adr/0001-inspection-transaction.md) for the inspection transaction and [ADR 0002](https://github.com/ghuntley/preflight/blob/main/docs/adr/0002-rust-workers-and-content-cache.md) for workers and caching.

## Policy

| mode | what happens |
| --- | --- |
| `redact` | Default. Replace text findings and sanitize supported attachments. Forward the rewritten request; return HTTP 409 when a finding has no safe replacement. |
| `no-go` | Return HTTP 409 for any non-allowlisted finding. Nothing reaches underclass. The response contains opaque finding IDs, never the secret. |
| `advisory` | Report findings and forward the original content. Useful for tuning; detected secrets can reach the model in this mode. |

All modes reject acquisition or extraction failures. Inspection failures return `422`, malformed JSON `400`, body limits `413`, and deadline exhaustion `408`. Preflight-generated errors use static codes rather than matched content.

## Configuration

Optional TOML configuration, supplied explicitly:

```
preflight serve --config /path/to/config.toml
```

| key | default | meaning |
| --- | --- | --- |
| `bind` | `127.0.0.1:8081` | listen address |
| `upstream` | `http://127.0.0.1:8080` | underclass base URL, without `/v1` |
| `mode` | `redact` | enforcement policy |
| `sandbox` | `true` | isolate attachment workers with Bubblewrap |
| `allow_page_redaction` | `false` | permit blacking out a PDF page when a finding cannot be mapped to an OCR region |
| `max_body_bytes` | `67108864` | maximum request body size: 64 MiB |
| `request_timeout_secs` | `180` | deadline covering inspection and waiting for upstream response headers |
| `carnet` | unset | path to a JSON array of exact secret SHA-256 hashes |
| `stopwords` | empty | exact whole-secret exemptions |
| `upstream_entropy` | `false` | apply upstream entropy thresholds to enabled rules |
| `cache_dir` | `$XDG_CACHE_HOME/preflight` or `~/.cache/preflight` | persistent verdicts and sanitized artifacts |
| `control_socket` | `/tmp/preflight-control.sock` | local cache administration socket |
| `resolver.file_api_base` | unset | trusted OpenAI-compatible file API base, such as `https://api.openai.com/v1` |

Example:

```
bind = "127.0.0.1:8081"
upstream = "http://127.0.0.1:8080"
mode = "redact"
carnet = "/run/secrets/preflight-carnet.json"

[cache]
memory_entries = 50000
disk_entries = 100000
artifact_bytes = 1073741824
max_artifact_bytes = 67108864
ttl_secs = 604800
```

Runtime environment:

| variable | purpose |
| --- | --- |
| `PREFLIGHT_CONFIG` | configuration path for `serve` |
| `PREFLIGHT_CLIENT_KEY` | optional bearer credential required from clients |
| `PREFLIGHT_UPSTREAM_KEY` | optional replacement credential sent to underclass |
| `PREFLIGHT_FILE_API_KEY` | separate credential for the configured file API |
| `PREFLIGHT_CONTROL` | socket path for cache administration commands |
| `RUST_LOG` | logging filter; output is structured JSON |

Send SIGHUP to reload configuration and the carnet. A failed reload keeps the previous runtime active, and admitted requests retain their original snapshots. Listener, cache directory/budgets, and control-socket changes require a restart. On NixOS: `systemctl reload preflight`.

## CLI

```
preflight serve [--config PATH]
preflight check [--config PATH]
preflight cache status [--socket PATH]
preflight cache purge [--scope SCOPE_ID] [--socket PATH]
```

`check` validates configuration and compiles the detection rules. Cache commands talk to the running daemon through a mode-`0600` Unix socket. On NixOS, run administration as root.

## Endpoints

| route | auth | purpose |
| --- | --- | --- |
| `POST /v1/responses` | client key, if configured | inspect a Responses request, then stream through underclass |
| `POST /v1/chat/completions` | client key, if configured | inspect a Chat Completions request, then stream through underclass |
| `GET /v1/models` | client key, if configured | forward model discovery |
| `GET /healthz` | none | local liveness probe |
| `GET /readyz` | none | readiness after startup toolchain checks |
| `GET /metrics` | none | aggregate Prometheus counters and request-to-headers timings |

Every handled response carries `x-request-id`, including blocked requests, authentication failures, model discovery, and unmatched routes. Completed inspections also attach `x-preflight-finding-count`. JSON logs correlate requests and safe finding IDs without recording prompts or matched secrets. Unknown routes are not pass-through routes.

### Correlation with underclass

One request gets one ID across both proxies:

```
client β†’ preflight β†’ underclass
         generates   preserves
         x-request-id: 9031b620-66aa-4a59-9228-bb7f40f567a1
```

Preflight creates a fresh UUIDv4 at ingress and sends it to underclass as `x-request-id`. Underclass validates and preserves it in its response, routing logs, retries, and request history. Both services use the log field `request_id`, so search that value in either service to follow the same request. Preflight replaces caller-supplied IDs; underclass generates its own when called directly with an absent, invalid, or duplicate ID.

Both services use `x-request-id` exclusively. IDs are diagnostic metadata, not authentication or proof of inspection. See [ADR 0005](https://github.com/ghuntley/preflight/blob/main/docs/adr/0005-cross-proxy-request-correlation.md).

## JSON logs

Preflight prints one JSON object per log line. These representative excerpts omit tracing's `span` and `spans` metadata for readability; in the full output, inspection events carry the request ID and inspection-profile fingerprint in their inspection span, nested under the request span. Timestamps, IDs, and timings below are illustrative.

The headings identify the configured policy. The current log schema does **not** include a `mode` or `action` field, and HTTP 200 alone does not distinguish redaction from advisory forwarding. Successful-forwarding examples assume underclass returns 200.

### No secret found β€” any mode

Inspection completes with zero findings, and the request continues normally:

```
{"timestamp":"2026-09-20T12:00:00.001Z","level":"INFO","fields":{"event":"inspection.completed","finding_count":0}}
{"timestamp":"2026-09-20T12:00:00.024Z","level":"INFO","fields":{"event":"request.policy_completed","request_id":"9031b620-66aa-4a59-9228-bb7f40f567a1","status":200,"duration_ms":24}}
```

There is no `inspection.finding` event for this request. An allowlisted fixture also contributes no finding.

### Secret found β€” `redact`

A text finding produces a warning with its rule and opaque finding ID. Preflight replaces the detected span with a token such as `[REDACTED:github-pat]`, then forwards the sanitized request:

```
{"timestamp":"2026-09-20T12:01:00.002Z","level":"INFO","fields":{"event":"inspection.completed","finding_count":1}}
{"timestamp":"2026-09-20T12:01:00.002Z","level":"WARN","fields":{"event":"inspection.finding","finding_id":"8a4f0571-4751-4d64-94da-55c3626a12de","rule_id":"github-pat"}}
{"timestamp":"2026-09-20T12:01:00.031Z","level":"INFO","fields":{"event":"request.policy_completed","request_id":"aa7445a8-b378-4b93-8ed1-b25ae80dfc01","status":200,"duration_ms":31}}
```

For rebuilt attachments, an additional `attachment.inspected` event includes `finding_count` and `rebuilt: true`. A finding that cannot be safely rewritten is blocked instead.

### Secret found β€” `no-go`

The finding is reported, then preflight returns 409 without sending the request to underclass:

```
{"timestamp":"2026-09-20T12:02:00.002Z","level":"INFO","fields":{"event":"inspection.completed","finding_count":1}}
{"timestamp":"2026-09-20T12:02:00.002Z","level":"WARN","fields":{"event":"inspection.finding","finding_id":"4b23f164-cc68-4
ghuntley10
🟧 echo.github ⭐Preflight is a local proxy that scans LLM requests and attachments for secrets, defaults to redaction, supports blocking and advisory modes,ghuntleyβ€”β€”
🟧 hnShow HN: Tokenhush – keeps your secrets out of what Claude Code sendsfregie30
🟠 redditI made a live globe of every server Claude Code talks to, and it blocks the ones you didn't allow
ClaudeAI
Elegant-Schedule819811
🟠 redditHow do you control what context your coding agent sends to an LLM? I built a local tool to measure and audit it β€” looking for blunt feedback
LocalLLaMA
1982_miguel07
🟠 redditClaude kept refusing to put my API key in a .env, so I made a proxy where it never sees the key
ClaudeAI
Otherwise_Ship_978205

Interpretation history

Decision trace