2026-10-11 16:37 UTC

ghuntley claims Underclass's released local proxy pools ChatGPT/Codex and GitHub Copilot subscriptions behind an OpenAI-compatible endpoint with persistent session affinity, automatic quota cooldowns, and fail-fast saturation, reducing account-management overhead and prompt-cache disruption for agent workloads.

state: watchingheat: lowuncertainty: mediumconvergesscott: mediumllm-routing api-proxy prompt-caching agent-infrastructureghuntley

What is this?

Underclass is a real, first-party Rust release by ghuntley (github.com/ghuntley/underclass, MIT): a localhost OpenAI-compatible proxy that pools multiple ChatGPT/Codex and GitHub Copilot subscriptions behind one endpoint, pinning each session to one account so the upstream prompt cache stays warm, cooling quota-exhausted accounts until their Retry-After window (default 1800s), and failing fast with the earliest reset when the pool is dry β€” device-flow OAuth, SQLite persistence, and mock-upstream e2e testing are all visible in the repo. The snippets do not establish production reliability under real multi-account load, measurable cache savings, adoption beyond the author, or the ToS/account risk of pooling paid subscriptions; those remain open. What the surrounding web does establish is that this is no longer a two-instance pattern: the llm-proxy/quota-management ecosystem now includes many independent subscription-pooling proxies (tokenproxy, chatgpt-codex-proxy, subctl-proxy, dario, mimo2api-hub, CLIProxyAPI descendants), several of which independently arrive at the same three mechanics β€” session/continuation affinity to one account, per-account 429 cooldown, automatic failover β€” suggesting subscription-pooling-with-affinity has become a recognized pattern for agent workloads.

Why it matters to Scott

The pooling-proxy pattern has spread from two instances to many, with independent builders converging on exactly the mechanics Scott's canon holds β€” cache-warm session pinning enacts Prefix-Caching Economics, the single OpenAI-compatible route matches his LiteLLM/company-gateway position, and 429 cooldown with failover matches his routing practice β€” a dated receipt that strengthens those arguments rather than merely illustrating them. Underclass's prompt_cache_key session affinity is also a concrete candidate mechanic (or alternative layer) for Ask's Codex-default LiteLLM path, but unproven production reliability, near-zero adoption, and ToS/account-ban exposure β€” salient given Scott's own OpenAI account-access dispute β€” cap actionability at medium.
ip:concept.prefix-caching-economicsip:concept.company-ai-gatewaydev:technology.litellmdev:project.askdev:concept.task-aware-model-routingradar:concept.llm-routingradar:concept.llm-gatewaysradar:concept.prompt-cachingradar:concept.usage-limitsradar:swobu-shareable-llm-switchboardradar:bounce-router-usage-failoverradar:codepress-subscription-cloud-agents
queries asked of Scott's wikis
  • prefix caching economics β€” session affinity and cache-warm routing for coding agents
  • LiteLLM gateway Codex default β€” routing, fallback, and quota cooldown behavior
  • subscription vs API-token cost model for agent workloads β€” pooling paid seats
  • local OpenAI-compatible proxy in front of providers β€” agent infrastructure setup
  • rate-limit and 429 handling in dev harnesses β€” account rotation and backoff
  • opencode connect β€” pointing coding agents at alternate model endpoints

Measured heat

now 0 pts/hpeak 3 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 509h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-20 12:22 (minted)⭐ origin echo-reconstructedUnderclass pools multiple ChatGPT/Codex and GitHub Copilot subscriptions behind one local OpenAI-compatible endpoint, pins sessions to accou
ghuntley on github (echo) Β· attributed from hn.story.49774704 Β· published time unknown
β€”
09-20 11:24first on hacker news Β· published Β· lag ?Underclass: An OpenAI-compatible pooling proxy that pins sessions to one account
ghuntley
β€”
09-20 11:24amplified on hacker newshn.story.49774704
ghuntley
peak 2 Β· 0 comments Β· 28% of case engagement
10-05 18:58amplified on hacker news πŸ‘‘hn.story.49968951
maxignol
peak 3 Β· 1 comments Β· 56% of case engagement
10-07 14:42amplified on hacker newshn.story.49993543
hugopuybareau
peak 1 Β· 0 comments Β· 15% of case engagement
09-20 12:20our radar first saw it Β· lag ?discovery anchor: hn.story.49774704β€”
pace: p23 vs 1032 stories at the 336h mark (now 509h old) β€” ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (4) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnUnderclass: An OpenAI-compatible pooling proxy that pins sessions to one account
Retrieved article excerpt

Open article Β· Retrieved 2026-09-20T12:21:58.513782+00:00

# underclass

**A local proxy that pools multiple ChatGPT/Codex and GitHub Copilot subscriptions behind one OpenAI-compatible endpoint.**

```
                       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                       β”‚            underclass              β”‚
                       β”‚                                    β”‚
 opencode ────────────►│  /v1/responses                     β”‚
 (any OpenAI-compatibleβ”‚  /v1/chat/completions              │────► chatgpt.com
  client)              β”‚  /v1/models                        β”‚      (N Codex subs,
                       β”‚                                    β”‚       OAuth device flow)
 web UI ◄─────────────►│  sticky sessions Β· health pool     β”‚
 (accounts, catalog,   β”‚  fail-fast saturation Β· tracing    │────► api.githubcopilot.com
  live request feed)   β”‚                                    β”‚      (M Copilot subs,
                       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜       GitHub device flow)
```

One subscription runs out of quota? It leaves rotation until its window resets β€” and comes back on its own. Sessions stay pinned to one subscription so upstream prompt caches stay warm. When *everything* is exhausted, the proxy fails fast with the earliest reset time instead of hanging.

---

## Why

Subscription-based model access has a per-account quota. One account is a ceiling; twenty accounts are a pool. underclass turns a pile of personal subscriptions into a single durable endpoint that behaves like one well-provisioned provider:

- **No client changes** β€” the surface is plain `/v1/*`; point opencode (or anything OpenAI-compatible) at it.
- **No quota whiplash** β€” exhausted accounts cool down and recover automatically; clients never see account churn.
- **No cache waste** β€” sessions are sticky, so the upstream prompt cache keeps working across turns.

## Quick start

```
UNDERCLASS_PROXY_KEY="$(openssl rand -hex 32)" \
UNDERCLASS_UI_TOKEN="$(openssl rand -hex 32)" \
cargo run -- serve
```

```
underclass listening on http://127.0.0.1:8080
web ui: http://127.0.0.1:8080/
```

Keep the generated values in a password manager or runtime secret file; underclass
never writes them to diagnostics. Open the web UI and paste the value supplied as
`UNDERCLASS_UI_TOKEN`.

1. Open the web UI and paste the configured admin token.
2. Click **Add account** β†’ pick *ChatGPT / Codex* or *GitHub Copilot* β†’ enter the device code at the shown URL. The account is labeled automatically with the account's email or username.
3. Repeat for every subscription you want in the pool.
4. Point opencode at the pool:

```
cargo run -- connect
opencode --provider underclass --model underclass/gpt-5.5
```

`underclass connect` writes the provider block and credentials into your global opencode config (`~/.config/opencode/opencode.json{,c}` + `auth.json`), idempotently and with backups. Re-run it any time; `--remove` undoes it.

## How routing works

- **Sticky sessions.** Requests carrying `prompt_cache_key` / `promptCacheKey` (opencode sends the session ID when configured with `setCacheKey: true`) always land on the same subscription. Bindings live for 24h, survive restarts, and rebind preferentially within the same backend when an account cools.
- **Health pool.** A quota response (`429`/usage-limit bodies) moves an account to *cooling* until `retry-after` (or a per-backend default). 401s trigger one token refresh + retry, then the account needs re-login. Cooling accounts stay configured and return to rotation automatically.
- **Fail fast.** If every account eligible for the requested model is cooling, the proxy answers `429` + `Retry-After` = earliest reset. No queuing.
- **Flat pool.** Codex and Copilot accounts compete by least-in-flight, filtered by per-backend model catalogs. Unknown model IDs pass through to Codex so new models work without proxy changes.
- **Pre-first-byte failover only.** Once a stream starts, upstream errors pass through β€” no silent re-send of half-finished turns.

Design decisions and their trade-offs live in [`docs/adr/`](https://github.com/ghuntley/underclass/blob/main/docs/adr) β€” start with [ADR 0010](https://github.com/ghuntley/underclass/blob/main/docs/adr/0010-unprefixed-routing-flat-pool.md) for the routing model.

## Configuration

Optional `~/.config/underclass/config.toml`:

| key | default | meaning |
| --- | --- | --- |
| `bind` | `127.0.0.1:8080` | listen address |
| `proxy_key` | minted on first run | bearer key clients must send to `/v1/*` |
| `ui_token` | minted on first run | admin token for the web UI + `/admin/api/*` |
| `codex_cooldown_secs` | `1800` | cooling window when upstream omits `retry-after` |
| `copilot_cooldown_secs` | `1800` | same, for Copilot |

Environment overrides: `UNDERCLASS_BIND`, `UNDERCLASS_PROXY_KEY`, `UNDERCLASS_UI_TOKEN`. For testing against a mock upstream: `UNDERCLASS_CODEX_UPSTREAM`, `UNDERCLASS_COPILOT_UPSTREAM` (default to the real endpoints).

State (credentials, sticky bindings, model catalog, minted keys) lives in `~/.local/share/underclass/pool.db`. Delete it to start fresh.

## CLI

```
underclass serve [--bind ADDR]
underclass connect [--base-url URL] [--api-key KEY] [--model MODEL]
                   [--project] [--no-default-model] [--dry-run] [--remove]
```

`connect` targets the global opencode config by default; `--project` writes `./.opencode/opencode.json` instead. `--dry-run` prints the merged documents without writing.

## Endpoints

| route | auth | purpose |
| --- | --- | --- |
| `POST /v1/responses` | proxy key | Responses API, streamed through to the pool |
| `POST /v1/chat/completions` | proxy key | Chat Completions, same |
| `GET /v1/models` | proxy key | union catalog with merged limits |
| `GET /` | none | web UI |
| `GET /admin/api/state` | admin token | accounts, catalog, last 200 requests |
| `POST /admin/api/flows` | admin token | start a device-flow onboarding |
| `GET /admin/api/flows/{id}` | admin token | poll an onboarding flow |
| `POST /admin/api/accounts/{id}/enable | disable | relogin` |
| `DELETE /admin/api/accounts/{id}` | admin token | remove from pool |
| `GET | PUT /admin/api/catalog/{backend}` | admin token |
| `GET /admin/api/client-key` | admin token | retrieve the proxy key for `connect` |

Every response carries `x-request-id`; logs are JSON (`RUST_LOG` filters, `--log-format json|pretty`) and each request logs the account (label) that served it.

## Models

The catalog is data, not code: seeded with the Codex families (`gpt-5.4`, `gpt-5.4-mini`, `gpt-5.3-codex-spark`, `gpt-5.5`, `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna`, `gpt-6-astra`) and Copilot's live `/models` list. Edit it in the UI or via the admin API; routing eligibility and the opencode model block follow it. See [ADR 0007](https://github.com/ghuntley/underclass/blob/main/docs/adr/0007-config-driven-model-catalog.md).

## Nix flake

The repository is a Nix flake: it exposes the CLI as a package/app and the devenv development shell as `devShells.default`.

Run the proxy without installing:

```
nix run github:ghuntley/underclass -- serve
```

Install it into your profile:

```
nix profile install github:ghuntley/underclass
```

Use the devenv shell (Rust toolchain, cargo) for development:

```
nix develop --no-pure-eval
cargo test
```

`--no-pure-eval` is required for the devenv shell (devenv inspects the working directory; this matches devenv's own flake template). The package and app outputs are pure β€” `nix run` and `nix profile install` need no flags.

Long-term devenv users can keep using `devenv shell` / `devenv test` directly β€” `nix develop` and `devenv shell` activate the same `devenv.nix`.

### NixOS module

Add underclass to your flake inputs and import its module:

```
{
  inputs.underclass.url = "github:ghuntley/underclass";

  outputs =
    { nixpkgs, underclass, ... }:
    {
      nixosConfigurations.my-host = nixpkgs.lib.nixosSystem {
        system = "x86_64-linux";
        modules = [
          underclass.nixosModules.default
          {
            services.underclass = {
              enable = true;
              bindAddress = "127.0.0.1:8080";
              environmentFile = "/run/secrets/underclass.env";
              settings = {
                codex_cooldown_secs = 1800;
                copilot_cooldown_secs = 1800;
              };
            };
          }
        ];
      };
    };
}
```

The runtime environment file can provide credentials without placing them in the Nix store:

```
UNDERCLASS_PROXY_KEY=sk-underclass-...
UNDERCLASS_UI_TOKEN=...
```

The service uses a dynamic user, persists its database in `/var/lib/underclass`, binds to localhost by default, and leaves the firewall closed. Set `services.underclass.openFirewall = true` only when intentionally binding beyond localhost.

The flake also exports `overlays.default`. Validate the module and its QEMU machine test with:

```
nix build .#checks.x86_64-linux.underclass-module
nix build .#checks.x86_64-linux.underclass-vm
```

Notes:

- The package builds from the committed `Cargo.lock`; dependency versions are pinned there.
- `nix build` skips `cargo test` because the property tests compile the Hegel engine as a build step, which needs network access that the Nix sandbox denies. CI runs the full suite via devenv (see `.github/workflows/ci.yml`).

## Security

- Access/refresh tokens, authorization headers, and prompt bodies are **never** logged.
- Account labels (email/username) appear in logs and the UI by design; raw UUID account IDs are truncated.
- `auth.json` written by `connect` uses `0600`. Never commit `pool.db` or `*.bak`.
- The proxy binds to localhost by default; put it behind a tunnel only if you understand the exposure.
- OAuth tokens rotate: every Codex refresh persists the new refresh token immediately.

## Development

```
cargo build
cargo test        # 43 unit tests + 8 Hegel property tests + e2e suite
```

Testing is two-tier: plain unit tests for exact behavior (headers, merges, redaction), and [Hegel](https://hegel.dev) property tests over the pure pool core β€” stickiness stability, health-state invariants, saturation minimums, TTL/cap bounds. The core (`src/pool.rs`, `src/health.rs`) is synchronous with an injected clock; async lives only at the edges. New backends implement the `provider::Backend` trait and register β€” nothing else changes.

Agent conventions and the ADR policy are in [`AGENTS.md`](https://github.com/ghuntley/underclass/blob/main/AGENTS.md). Architecture decision records: [`docs/adr/`](https://github.com/ghuntley/underclass/blob/main/docs/adr).

## License

[MIT](https://github.com/ghuntley/underclass/blob/main/LICENSE)

## Project layout

```
src/
  main.rs      bootstrap, router, background tasks
  config.rs    TOML + env config
  models.rs    domain types (Account, status, catalog, log entries)
  store.rs     SQLite persistence (accounts, bindings, catalog, config)
  pool.rs      pure pool core: stickiness, eligibility, selection
  health.rs    quota classification, retry-after parsing
  provider.rs  Backend trait
  codex.rs     ChatGPT/Codex backend (device flow, refresh, headers)
  copilot.rs   GitHub Copilot backend (device flow, catalog, headers)
  tokens.rs    single-flight token refresh
  proxy.rs     /v1/* handlers, failover, streaming
  logging.rs   structured logs, correlation IDs, redaction
  ui.rs        admin API
  ui.html      embedded single-page web UI
  cli.rs       `connect` verb, JSONC-safe opencode config merge
tests/
  properties.rs  Hegel property tests over the pure core
  e2e.rs         full-pool story against a mock upstream
```
ghuntley20
🟧 echo.github ⭐Underclass pools multiple ChatGPT/Codex and GitHub Copilot subscriptions behind one local OpenAI-compatible endpoint, pins sessions to accoughuntleyβ€”β€”
🟧 hnShow HN: Codex and Claude subscriptions pooled together with whoever you trustmaxignol31
🟧 hnShow HN: Coding easily with several Claude and Codex accountshugopuybareau10

Interpretation history

Decision trace