2026-10-11 17:15 UTC

ggeorgovassilis releases llm-gauze, an OpenAI-compatible HTTP gateway that detects and remediates open-weight LLM quirks (malformed tags, empty responses, stuck loops, context overflows) before clients see them, with logging, metrics, and Docker deployment.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumagent-harnesses local-inference model-gatewayggeorgovassilis

What is this?

The case describes llm-gauze, an OpenAI-compatible HTTP gateway by ggeorgovassilis that sits in front of open-weight LLMs to detect and remediate common quirks โ€” malformed tags, empty responses, stuck generation loops, context overflows โ€” before clients see them, with logging, metrics, and Docker deployment. The web search results returned general LLM gateway comparisons (LiteLLM, LLM Gateway, AegisGate, etc.) but did not surface the llm-gauze project itself; the only direct signal is the evidence title referencing a 'Show HN: Gauze fixes (some) open-weight LLM deficiencies' post. Without the Show HN thread or the project's repo/docs in the snippets, the exact feature set, maturity, and adoption signals cannot be verified from the supplied material.

Why it matters to Scott

ggeorgovassilis's llm-gauze independently arrives at the gateway-layer remediation pattern Scott's architecture prescribes (architectural-containment, router-as-kernel) and his own stack partially implements (LiteLLM proxy, ask's multi-format tool-call parsing with json_repair). The specific quirk coverage โ€” stuck generation loops, empty responses, context overflow remediation at the gateway โ€” extends beyond his current LiteLLM+ask stack, which handles malformed tool calls but not runtime generation pathologies. This is a concrete artifact Scott would evaluate for adoption or reference, not merely another illustration of his pattern.
dev:technology.litellmdev:concept.multi-format-tool-call-parsingip:concept.architectural-containmentip:concept.verification-loopsip:concept.answer-failure-classesdev:project.askdev:concept.agent-authored-context-compactionip:concept.non-determinismradar:abliterated-weights-agent-backdoorradar:aa-agentperf-local-benchmarkradar:49ide-spatial-agent-workspaceradar:adaptive-kv-cache-streamingradar:agentgauntlet-failure-benchmarkradar:agentic-context-management-paper
queries asked of Scott's wikis
  • model gateway pattern for local/open-weight inference reliability
  • agent harness infrastructure: handling model quirks at the gateway layer
  • open-weight LLM failure modes: malformed tags, empty responses, stuck loops, context overflow
  • local inference deployment: Docker-based gateway proxies for agent workflows
  • observability and metrics for model gateway remediation actions

Measured heat

now 0 pts/hpeak 1 pts/hcomments 0/hpeers p16momentum: steady2 platformsage 434h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-23 14:00โญ origin echo-reconstructedOriginal project repo, authored and announced by the same person who posted to HN. README: "llm-gauze is an HTTP gateway that sits in front
ggeorgovassilis (George Georgovassilis) on github (echo) ยท attributed from hn.story.50005230
โ€”
10-08 12:51first on hacker news ยท published ยท +358.9hShow HN: Gauze fixes (some) open-weight LLM deficiencies
ggeorgovassilis
โ€”
10-08 12:51amplified on hacker news ๐Ÿ‘‘hn.story.50005230
ggeorgovassilis
peak 1 ยท 1 comments ยท 49% of case engagement
10-10 15:37amplified on hacker newshn.story.50033976
ggeorgovassilis
peak 1 ยท 1 comments ยท 49% of case engagement
10-08 15:39our radar first saw it ยท +361.7hdiscovery anchor: hn.story.50005230โ€”

Evidence (3) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnShow HN: Gauze fixes (some) open-weight LLM deficiencies
Retrieved article excerpt

Open article ยท Retrieved 2026-10-08T17:49:33.819540+00:00

# llm-gauze

[CI](https://github.com/ggeorgovassilis/llm-gauze/actions/workflows/ci.yml)

llm-gauze is an HTTP gateway that sits in front of a local LLM (served via an
OpenAI-compatible API) and works around its shortcomings: transient errors
without retries, silently hung or looping models, empty or sloppy responses,
context-window overflows, runaway reasoning, and malformed tool calls. It
logs every exchange and remediates what it can before the client ever sees
it.

## Quick start

llm-gauze ships as a published container image and runs with Docker Compose.

1. Create your configuration:

   ```
   cp .env.example .env
   ```
2. Point `LLM_BASE_URL` in `.env` at your local LLM's OpenAI-compatible
   endpoint (default `http://host.docker.internal:14434`, which reaches a
   host-side LLM on port 14434 from inside the container).
3. Start the gateway:

   ```
   docker compose up
   ```

   This runs the published image `ghcr.io/ggeorgovassilis/llm-gauze:latest`
   (defined in `docker-compose.yml`). To pin a specific release, override the
   tag โ€” for example `ghcr.io/ggeorgovassilis/llm-gauze:7`.

The gateway listens on `http://localhost:9317` and exposes an
OpenAI-compatible API (e.g. `POST /v1/chat/completions`), forwarding to the
`LLM_BASE_URL` in `.env`.

### File ownership

The container runs as your host user (`UID`/`GID`, default `1000`) so the
files it writes to the bind-mounted `data/` directory are owned by you, not
root. If a `data/` directory was created earlier as root, fix ownership once
before the non-root container can write to it:

```
sudo chown -R "$(id -u):$(id -g)" data
```

## Health check

```
curl http://localhost:9317/health
```

Returns `{"status": "ok", "upstream": "<LLM_BASE_URL>"}` when the gateway is up.
Operational metrics are exposed at `/metrics` (Prometheus text format, or JSON
with `Accept: application/json`).

## Configuration

Every setting is documented in [`docs/configuration.md`](https://github.com/ggeorgovassilis/llm-gauze/blob/main/docs/configuration.md).

## How it works

llm-gauze sits between your client and the local LLM, recording every exchange:

```
flowchart LR
    Client[Your client] -->|OpenAI-compatible API| Gauze[llm-gauze]
    Gauze -->|forwards| LLM[Local LLM]
    Gauze -.->|logs every exchange| Store[(data/*.jsonl)]
```

 Loading

When the model misbehaves, llm-gauze detects it and remediates what it can before
you ever see it โ€” retrying transient failures, nudging empty replies, cleaning
leaked thinking tags, breaking loops with varied sampling:

```
sequenceDiagram
    participant C as Client
    participant B as llm-gauze
    participant L as Local LLM

    C->>B: POST /v1/chat/completions
    B->>L: forward
    L-->>B: error, hang, loop, or sloppy reply
    B->>B: detect & classify
    B->>L: remediate (retry / nudge / repair)
    L-->>B: clean completion
    B-->>C: chat.completion
```

 Loading

See [`docs/architecture.md`](https://github.com/ggeorgovassilis/llm-gauze/blob/main/docs/architecture.md) for the full design and
[`docs/configuration.md`](https://github.com/ggeorgovassilis/llm-gauze/blob/main/docs/configuration.md) for every setting.

## Developing

See [`DEVELOPING.md`](https://github.com/ggeorgovassilis/llm-gauze/blob/main/DEVELOPING.md) for local development, running the tests,
and adding a detector or remediation.
ggeorgovassilis11
๐ŸŸง echo.github โญOriginal project repo, authored and announced by the same person who posted to HN. README: "llm-gauze is an HTTP gateway that sits in front ggeorgovassilis (George Georgovassilis)โ€”โ€”
๐ŸŸง hnGauze Corrects LLM Responsesggeorgovassilis11

Interpretation history

Decision trace