2026-10-11 16:38 UTC

MiaAI-Lab claims its released serving kit automatically selects EXL3 quantizations and serves Qwen3.8-27B through an OpenAI-compatible endpoint on a single 16GB NVIDIA GPU, potentially simplifying low-memory local deployment on Windows and Linux.

state: watchingheat: mediumuncertainty: highconvergesscott: mediumlocal-inference qwen model-quantizationMiaAI-Labturboderp

What is this?

MiaAI-Lab has a GitHub deployment kit containing a launcher, server, and configuration for serving an EXL3-quantized Qwen3.8-27B, with model weights hosted separately on Hugging Face. A social-post snippet credits the lab with a custom installer intended to make local use on a 16GB NVIDIA GPU easier. The supplied snippets do not establish automatic VRAM-based weight selection, Windows/Linux support, or an OpenAI-compatible endpoint for this kit; its repository name says 5.0bpw while its description says 3.5bpw, leaving the case’s claimed custom 2.0-bpw configuration unverified. No supplied snippet establishes turboderp’s involvement or independently measures this kit’s performance and quality on 16GB hardware.

Why it matters to Scott

MiaAI-Lab’s released quantized serving kit converges with Scott’s hardware-aware local-inference practice and offers a concrete runtime candidate to evaluate alongside gamepc’s Ollama service; the radar’s Qwen3.8 16GB quantization benchmark page tracks a related deployment choice, not this kit. Evaluation—not adoption—is warranted: automatic VRAM selection, OpenAI-compatible integration, Windows/Linux support, and useful 16GB performance remain unverified, with conflicting bit-width descriptions.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaradar:qwen38-27b-16gb-quant-benchmarkradar:concept.local-inferenceradar:concept.quantizationradar:concept.llm-serving
queries asked of Scott's wikis
  • local inference economics consumer GPU VRAM budgets
  • quantization quality tradeoffs context memory coding agents
  • OpenAI-compatible local endpoints agent harness integration
  • reproducible inference deployment automated runtime setup Windows Linux
  • self-hosted models cloud dependency model sovereignty

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 572h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-17 20:35 (minted)⭐ origin echo-reconstructedThe serving kit installs an isolated environment, selects weights by VRAM, and exposes an OpenAI-compatible endpoint, using a custom 2.0-bpw
MiaAI-Lab on github (echo) · attributed from hn.story.49745610 · published time unknown
—
09-17 19:46first on hacker news · published · lag ?Run QWEN3.8 27B on 16gb Nvidia GPUs
Pragmata
—
09-17 19:46amplified on hacker news 👑hn.story.49745610
Pragmata
peak 28 · 10 comments · 100% of case engagement
09-17 20:21our radar first saw it · lag ?discovery anchor: hn.story.49745610—
pace: p58 vs 1032 stories at the 336h mark (now 572h old) — ahead of agent-iap-credential-brokering (1.0x), behind anthropic-fourth-cyber-incident-review-miss (0.9x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnRun QWEN3.8 27B on 16gb Nvidia GPUs
Retrieved article excerpt

Open article · Retrieved 2026-09-17T20:23:06.555949+00:00

# Qwen3.8-27B on 16-32 GB Nvidia GPUs one-click install for Windows / Linux

[Qwen3.8-27B one-click install](https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install/blob/main/assets/intro.png)

by [Mia'a AI Lab](https://x.com/MiaAI_lab)
  
  
[Sponsor me on GitHub](https://github.com/sponsors/MiaAI-Lab)
[Follow Mia on X](https://x.com/MiaAI_lab)

A serving kit for [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) in
**[turboderp](https://huggingface.co/turboderp)**'s EXL3 quants, on one consumer
NVIDIA card. It picks a quant that fits the card it finds, installs its own Python
environment, downloads the weights, serves an **OpenAI-compatible** endpoint, and
opens a chat UI. Windows and Linux, same behaviour.

It started as a 16 GB recipe — the 2.0 bpw quant is still that floor, and still the
one thing here that is not turboderp's own upload ([Mia-AiLab/Qwen3.8-27B-EXL3-2.0bpw](https://huggingface.co/Mia-AiLab/Qwen3.8-27B-EXL3-2.0bpw),
`SC_2.00bpw_H3_V3`). Everything from 2.5 bpw up is pulled from
[turboderp/Qwen3.8-27B-exl3](https://huggingface.co/turboderp/Qwen3.8-27B-exl3)
by revision. Which one you get is [decided by your VRAM](https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install#what-the-launcher-picks-for-your-gpu),
at setup, and you can change it any time.

> **Jump to your guide:** [Windows](https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install#windows) · [Linux](https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install#linux)

---

## Before you start (both systems)

|  |  |
| --- | --- |
| GPU | NVIDIA, 12 GB VRAM or more, compute capability 7.5+ (Turing and newer). 16 GB is the size this kit was built around. |
| Driver | 570 or newer (the default PyTorch build is cu128). |
| Python | **3.11 or newer, 64-bit.** The only thing you install by hand. |
| Node | **22.19+** from [nodejs.org](https://nodejs.org/) (current dsh). Older LTS (20) warns `EBADENGINE` and the chat UI may fail. Without Node, `/v1` still serves; the launcher says what is missing. |
| Disk | 9.7–22.9 GB per quant (see the table below), plus several GB for the Python environment and PyTorch. |

**Not needed:** CUDA Toolkit, Visual Studio Build Tools, Git. The engine arrives as a
[prebuilt wheel](https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install#prebuilt-wheels-no-compiler-needed); compiling is the fallback for
platforms no wheel covers.

Everything the kit installs stays inside its own folder — `.venv/`, `models/`,
`logs/`, `apps/`, `.dsh/`. Nothing goes into the system Python and nothing needs
administrator rights.

---

## Windows

Everything is in the `windows\` folder. Run the `.bat` files from Explorer
(double-click) or from a `cmd` window opened in the kit folder.

### 1. Install — `windows\START-HERE.bat`

Double-click it once. It opens a page in your browser and does the whole install
there: it shows what it found on the card, offers the model sizes that fit it, then
installs and downloads with a progress bar and a live log. Nothing is asked in the
console.

When the download finishes it **loads that model** and hands the page over to the
chat, so one double-click takes you from nothing to a working chat window.

```
windows\START-HERE.bat              install, then start what was installed
windows\START-HERE.bat --no-start   install only — for fetching a second size
```

Notes:

- The download is **resumable**. Closing the window, losing the connection or
  rebooting costs you nothing — it picks up from the byte it stopped at. A model
  left half-downloaded is shown as such in the menus, and running
  `windows\START-HERE.bat` again finishes it. Nothing offers to *start* a model
  until every weight file is on disk.
- Prefer the old console questions to the web page? Set `SETUP=console` in `.env`.

### 2. Every day after that — `windows\start.bat`

```
windows\start.bat          start a model that is already here
windows\start.bat setup    go to setup instead (same as START-HERE.bat)
```

It never downloads anything. What it does:

1. **Asks which model**, if more than one size is on disk. Enter takes the one that
   ran last, and it starts on its own after 45 seconds so an unattended machine
   still comes up.
2. **Checks free VRAM** right before the load — it wants the `GPU_MEM_GB` budget
   from `.env` plus a little margin. If that much is not free it lists the programs
   holding VRAM (browsers, games, Discord, other AI tools) and waits: Enter
   re-checks, `c` continues anyway, `q` quits, and it continues on its own after
   120 seconds. **Take this seriously on Windows** — with too little free VRAM the
   driver pages the model into system RAM instead of failing, and it then runs many
   times slower.
3. **Loads the model** and opens the [chat UI](https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install#the-deepseek-harness) at
   `http://127.0.0.1:3080/`.

If nothing is installed yet, or nothing finished downloading, it says so and offers
to run setup for you — double-clicking the wrong one is never a dead end.

Two files rather than one because they answer two different questions:
`windows\start.bat` never downloads, and `windows\START-HERE.bat --no-start` never
loads.

### 3. While it runs

- The console window it opened **is** the server. Closing it stops the model.
- Simplex puts an icon in the notification area — right-click for **Open Simplex**,
  **Restart the model**, **Show the Simplex folder**, **View the log** and
  **Quit Simplex**.
  (`TRAY=no` in `.env` turns it off.)
- Every launch writes a full transcript to `logs\`, so a crash that scrolls past is
  still readable afterwards.
- The first successful launch adds Start-menu and desktop shortcuts
  (`SHORTCUTS=no` in `.env` to skip that).

### 4. Stopping — `windows\stop.bat`

```
windows\stop.bat                 stop both the model and the chat UI
windows\stop.bat --harness-only  leave the model loaded, close the UI
windows\stop.bat --server-only   leave the UI running, unload the model
```

### 5. Windows troubleshooting

| symptom | what to do |
| --- | --- |
| "Simplex needs Python and cannot find it" | Install 64-bit Python 3.11+ from [python.org](https://www.python.org/downloads/) and tick **Add python.exe to PATH**, then run the file again. |
| Anything else | `windows\simplex.bat doctor` — see [below](https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install#simplex-doctor). |
| The model loads but crawls | Free VRAM (the check above told you what is holding it), or lower `CONTEXT_SIZE` / `GPU_MEM_GB` in `.env`. |
| "Images: off" in the Ready box | The vision tower did not fit next to your context. Lower `CONTEXT_SIZE` and restart, or pick a smaller quant. |
| The window closed and you missed the error | It is in `logs\` — newest file. `windows\simplex.bat logs` prints the tail. |
| You want to start completely over | `reset_new_user.bat` in the kit root deletes the weights, the venv, `.env`, the logs and the shortcuts, and keeps every tracked file. It asks you to type `RESET` first. |



---

## Linux

Everything is in the `linux/` folder. Run the scripts from the **kit root**; they
find their own way regardless of where you call them from.

If the files arrived without their execute bit (a zip, a copy off Windows), run them
as `bash linux/setup.sh` instead of `./linux/setup.sh`, or `chmod +x linux/*.sh linux/simplex` once.

### 1. Install — `./linux/setup.sh`

```
./linux/setup.sh
```

It creates `.env` from `.env.example` on the first run, asks the profile questions
**in the terminal** (`tools/profiles.py`) — there is no setup page on Linux — builds
`.venv`, installs PyTorch and the engine, downloads the weights, and stops. It does
not load a model.

The download is resumable: interrupt it and run `./linux/setup.sh` again to carry on.

### 2. Every day after that — `./linux/start.sh`

```
./linux/start.sh                 pick a downloaded model and serve it
./linux/start.sh --no-harness    serve /v1 only, no chat UI
./linux/start.sh -b              run in the background, output in logs/
./linux/start.sh --status        is a backgrounded one running?
```

It lists the models that finished downloading and asks which one (Enter is the one
used last; it auto-picks after 45 seconds), then serves:

```
http://localhost:8888/v1     the OpenAI-compatible API
http://127.0.0.1:3080/       the chat UI
```

`-b` is the honest equivalent of the Windows tray: it detaches, writes to
`logs/simplex-*.log`, and tells you where that log is and how to stop it. First-run
setup and the model menu still happen — written to the log instead of the screen.

**On a box with no desktop session** `webbrowser` has nothing to open, so the chat
address is printed for you to copy. Take the whole thing, **token and all** — see
[the note on the token](https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install#the-first-address-is-not-the-plain-one).

There is **no free-VRAM preflight on Linux** (that check is Windows-specific,
because Windows silently spills to system RAM instead of failing). If a load fails
with `Insufficient VRAM in split for model and cache`, lower `CONTEXT_SIZE` or
`GPU_MEM_GB` in `.env`, or close what is holding the card.

No tray icon and no desktop shortcuts either — those are Windows.

### 3. Stopping — `./linux/stop.sh`

```
./linux/stop.sh                 stop both
./linux/stop.sh --harness-only  leave the model loaded
./linux/stop.sh --server-only   leave the chat UI running
```

### 4. Linux troubleshooting

| symptom | what to do |
| --- | --- |
| `bash: ./linux/start.sh: Permission denied` | `chmod +x linux/*.sh linux/simplex`, or call it as `bash linux/start.sh`. |
| `$'\r': command not found` | The checkout has CRLF endings. `.gitattributes` prevents this; re-clone, or `sed -i 's/\r$//' linux/*.sh linux/simplex`. |
| Anything else | `./linux/simplex doctor` — see [below](https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install#simplex-doctor). |
| `Insufficient VRAM in split for model and cache` | Lower `CONTEXT_SIZE` or `GPU_MEM_GB` in `.env`, or run `./linux/simplex setup` and pick a smaller quant. |
| It compiled the engine for 20 minutes | No prebuilt wheel matched your CUDA line, torch version or Python. See [Prebuilt wheels](https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install#prebuilt-wheels-no-compiler-needed). |
| aarch64 / GB10 | No prebuilt engine wheel exists on any CUDA line, so it compiles. The script keeps cu130 there and sets `TORCH_CUDA_ARCH_LIST=12.0;12.1` for you. |



---

## One command, both systems

The files above are the double-click doors. Every verb, on either system, is
`simplex` — the same program (`tools/cli.py`), so the two cannot drift apart:

| Linux | Windows |
| --- | --- |
| `./linux/simplex <verb>` | `windows\simplex.bat <verb>` |

```
simplex setup                   install the environment and fetch a model
simplex start                   load a model and serve it
simplex start --no-harness      ...serving /v1 only
simplex start -b                ...in the background, log in logs/
simplex start -p 9000           ...on another port, just this once
simplex stop                    stop both
simplex stop --harness-only     ...and detach the UI, model still loaded
simplex restart                 stop, then start
simplex status                  what is running, which model, which ports
simplex status --json           the same, for scripts
simplex logs -f                 follow the launcher log
simplex models                  what is on disk, and what is half-downloaded
simplex harness start           attach the UI to a server already running
simplex harness stop|status|open|settings
simplex doctor                  check this machine before blaming the model
```

`simplex` with no verb prints the help and then the status. Every verb takes
`--help`. It is not on your 
Pragmata2810
🟧 echo.github ⭐The serving kit installs an isolated environment, selects weights by VRAM, and exposes an OpenAI-compatible endpoint, using a custom 2.0-bpwMiaAI-Lab——

Interpretation history

Decision trace