Retrieved article excerpt
Open article ยท Retrieved 2026-09-25T16:30:16.172096+00:00
# llama.cpp: fork with multi gpu acceleration even for models bigger than total gpu mem
> **Big thanks to [@csantiago78](https://github.com/csantiago78)**: the expert cache
> here builds on their implementation in llama.cpp PR
> [#27861](https://github.com/ggml-org/llama.cpp/pull/27861) ("GPU-resident LRU cache
> for host-offloaded MoE expert weights"), the first to get a working hot-expert
> cache into llama.cpp. This fork takes it further: filling all free VRAM, smarter
> eviction, CPU/GPU overlap and fused kernels.
For MoE models much larger than VRAM: every expert stays in system RAM, and all
VRAM left after the KV cache becomes a live cache of the experts actually being
used. The GPUs compute cached experts while the CPU computes the rest, in parallel.
**Results, 2x RTX 3090 (48 GB) + 125 GB RAM, single stream, temp 0, `-c 1024`:**
| model | size | stock t/s | fork t/s | gain | PPL (stock -> fork) |
| --- | --- | --- | --- | --- | --- |
| GLM-5.3-Flash 3.0-bit, original GGUF | 117 GB | 12.3 | ~25 | 2.0x | 3.5534 -> 3.5534 |
| GLM-5.3-Flash 3.0-bit, [Q4\_K attention GGUF](https://huggingface.co/neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF) | 106 GB | 12.3 | **27.72** | 2.3x | 3.5534 -> 3.5871 (+0.95%) |
| MiMo-2.6-Flash-RL IQ3\_XXS | 132 GB | 4.7 | **9.9** | 2.1x | unchanged (no requant) |
| Qwen3.8-Flash-Next UD-IQ4\_XS | 88 GB | 29.7 | **32.1** | 1.08x | unchanged (no requant) |
GLM prompt: "generate smallest html tetris game."; MiMo/Qwen prompt: "write smallest
html tetris game" (both temp 0). PPL: wikitext-2, 40 x 512-token chunks (GLM only).
MiMo needed `-fitt 8000` on both stock and fork to avoid autofit OOM-ing on this
arch/quant combo; the others loaded fine with default fit.
Qwen's gain is small because stock's autofit already placed most of its experts on
GPU by default here โ little room left for the cache to improve on. The big wins
(GLM, MiMo) are on models where default placement leaves most expert work on the CPU.
Models:
- GLM original GGUF (tested): [pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF](https://huggingface.co/pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF),
the 3.0-bit file. It works as-is with this fork, no conversion needed, and it's the
quality reference (unchanged perplexity). Most of the speedup comes from the fork,
not the requantization.
- GLM Q4\_K attention variant (same experts, non-expert Q8\_0 weights requantized to Q4\_K):
[neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF](https://huggingface.co/neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF).
**Same VRAM, different use.** Stock llama.cpp and this fork get the same 48 GB; what
differs is what it holds. Each token uses only 8 of the 288 experts in each layer.
- Stock places experts statically, whole layers at a time: ~36 GB fits all 288
experts of ~14 of the 42 MoE layers. Most of that VRAM holds experts the current
token doesn't touch, so only ~33% of each token's expert work runs on GPU and the
CPU does ~67%, one after the other.
- This fork fills the same VRAM with the ~100 most-used experts of every layer.
Usage is skewed, so those cover ~85% of what tokens actually pick: ~85% of expert
work runs on GPU and the CPU does ~15%, at the same time as the GPUs.
| | expert work on GPU | expert work on CPU | decode t/s |
| --- | --- | --- | --- |
| stock (static whole layers) | ~33% | ~67% | 12.3 |
| this fork (cache of hot experts) | ~85% | ~15%, in parallel | 26-28 |
So stock can't reach 2x on the same hardware: without an expert cache, extra VRAM
mostly holds experts that aren't being used.
Run (with the faster Q4\_K attention GGUF from [neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF](https://huggingface.co/neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF)):
```
llama-server -m GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf \
-np 1 -c 1024 -t 6 --cpu-moe -nr --moe-expert-cache -1
```
- `--cpu-moe` keeps all experts in RAM; `--moe-expert-cache -1` sizes the cache
per GPU from the VRAM free after KV and compute buffers. Set `-c` explicitly:
without it autofit grows the context and takes the VRAM the cache needs.
- `-t 6` suits an 8-core CPU (leave cores to drive the GPUs).
- Cache hit rate is ~85% after warm-up. `LLAMA_MOE_CACHE_STATS=1` logs it.
- Tuning: `LLAMA_MOE_CACHE_POLICY` (`add` default, `halve`, `window`),
`LLAMA_MOE_CACHE_MARGIN_MB` (VRAM left free, default 1024),
`LLAMA_MOE_CACHE_SWAP_FRAC` (share of token time for uploads, default 0.25).
- `GGML_SCHED_PROF=1` prints where each token's host time goes.
**More GPUs (estimate, only 2 tested).** Nothing assumes two GPUs: each GPU gets
its own cache, sized from its free VRAM. Each extra GPU adds cache room, so more
of every token's experts are hits and less work falls to the CPU. For this model,
per token today: ~20 ms GPU work on non-expert layers, ~10 ms CPU on missed experts.
| GPUs (24 GB each) | cache room | slots/layer (of 288) | hit rate | CPU miss time | decode t/s |
| --- | --- | --- | --- | --- | --- |
| 2 (measured) | ~34 GB | ~100 | 85-88% | ~10 ms | 26-28 |
| 3 | ~58 GB | ~170 | ~95% | ~3-4 ms | ~33-35 |
| 4 | ~82 GB | ~240 | ~99% | ~1 ms | ~38-42 |
| 5+ | whole model | 288 | 100% | 0 | ~40-45 (plateau) |
The plateau is the ~20 ms GPU part: with the default layer split each layer runs on
one GPU at a time, so extra GPUs add cache room, not speed on that part. System RAM
must still hold all experts. Cards in x4 PCIe slots upload experts slower, so the
cache warms up slower. Reports from 3+ GPU setups are welcome.
What's in it: the GPU expert cache from PR [#27861](https://github.com/ggml-org/llama.cpp/pull/27861)
(csantiago78), extended with VRAM-filling auto-sizing, usage-driven eviction that
only swaps when the upload pays back, CPU/GPU overlap per layer, scheduler barrier
fixes and fused gate kernels. GLM-5.3-Flash support comes from PRs
[#27773](https://github.com/ggml-org/llama.cpp/pull/27773) and
[#27917](https://github.com/ggml-org/llama.cpp/pull/27917) (timkhronos); stock
llama.cpp can't load GLM-5.3-Flash yet.
**Prefill warm start**, implemented independently for this fork: the cache observes
which experts the prompt itself selects during prefill and preloads them before the
first generated token, instead of starting cold and only learning from decode.
[@sdroege](https://github.com/sdroege) explored the same idea independently
in the [PR #27861 discussion](https://github.com/ggml-org/llama.cpp/pull/27861)
with their own patch; worth checking out too.
**Other notable tweaks:**
- **Swap budget from measured cost, not a guess**: each step measures real upload
time (ms/expert) and real token time, then computes how many swaps fit in
`LLAMA_MOE_CACHE_SWAP_FRAC` of a token (default 25%) โ instead of a fixed
swaps-per-step constant.
- **Pay-back filter on every eviction**: a swap only happens if the candidate's
measured usage beats the cached victim's by more than what the upload itself
costs in CPU-equivalent time, so churn can't cost more than it saves.
- **Usage tracking with decay** (`LLAMA_MOE_CACHE_POLICY`: `add` default, `halve`,
`window`): recent use counts more than old use, so the cache follows shifts in
which experts are hot instead of freezing on early-token bias.
- **CPU/GPU overlap inside a layer**: the GPU cache chain is queued and its inputs
copied *before* the CPU miss chain runs, so both compute at the same time instead
of the scheduler serializing them.
- **Scheduler fixes upstream benefits from too**: no host barrier between two GPU
splits when neither reads host memory, and a new split is inserted exactly when
another GPU's result is needed mid-split โ both apply to any multi-GPU llama.cpp
workload, not just this cache.
- **Fused CUDA kernel** for the hyper-connection/KDA gate chain
(`MUL -> ADD|SCALE -> SIGMOID -> SCALE`, one kernel instead of four), and GLM5-Next's
KDA Q/K norm collapsed from `rms_norm`+`scale` into one `l2_norm` op.
- **Repacked CPU experts stay cacheable**: uploads read raw bytes straight from the
GGUF file (recorded per-tensor file offsets) when host memory holds a
repack-transformed layout instead of the on-disk one.
- `GGML_SCHED_PROF=1` and `LLAMA_MOE_CACHE_STATS=1` for live profiling: barrier vs.
copy wait time, fill %, in-flight uploads, queue depth, hit rate.
**About the author of this fork**: I'm actively looking for an AI engineering/research
role and open to relocating out of Eastern Europe. If this work is useful to you or
your team, reach out: [linkedin.com/in/neuralll](https://www.linkedin.com/in/neuralll/)
---
[llama](https://raw.githubusercontent.com/ggml-org/llama.brand/refs/heads/master/cover/llama-cpp/cover-llama-cpp-dark.svg)
**LLM inference in C/C++**
[License: MIT](https://opensource.org/licenses/MIT)
[Release](https://github.com/ggml-org/llama.cpp/releases?q=tag:v0)
[Nightly](https://github.com/ggml-org/llama.cpp/releases?q=b)
[Server](https://github.com/ggml-org/llama.cpp/actions/workflows/server.yml)
[Docker](https://github.com/ggml-org/llama.cpp/actions/workflows/docker.yml)
[Winget](https://github.com/ggml-org/llama.cpp/actions/workflows/winget.yml)
[ggml](https://github.com/ggml-org/ggml) / [ops](https://github.com/ggml-org/llama.cpp/blob/master/docs/ops.md) / [maintainer PRs](https://github.com/ggml-org/llama.cpp/issues?q=is%3Apr%20is%3Aopen%20draft%3AFalse%20(author%3Argerganov%20OR%20author%3AKitaitiMakoto%20OR%20author%3Adanbev%20OR%20author%3Aaldehir%20OR%20author%3Amax-krasnyansky%20OR%20author%3ACISC%20OR%20author%3Aggerganov%20OR%20author%3Aam17an%20OR%20author%3Ajhen0409%20OR%20author%3Abartowski1182%20OR%20author%3Anikwen%20OR%20author%3Ahipudding%20OR%20author%3Aravi9%20OR%20author%3AServeurpersoCom%20OR%20author%3Apwilkin%20OR%20author%3Areeselevine%20OR%20author%3Angxson%20OR%20author%3Ajeffbolznv%20OR%20author%3Amarty1885%20OR%20author%3A0cc4m%20OR%20author%3ATitaniumtown%20OR%20author%3Aangt%20OR%20author%3AIMbackK%20OR%20author%3Aarthw%20OR%20author%3AJohannesGaessler%20OR%20author%3AORippler%20OR%20author%3Aruixiang63%20OR%20author%3Axctan%20OR%20author%3Aallozaur%20OR%20author%3Ayomaytk%20OR%20author%3Aaendk%20OR%20author%3Awine99%20OR%20author%3Agaugarg-nv%20OR%20author%3Ataronaeo%20OR%20author%3Aforforever73%20OR%20author%3Alhez%20OR%20author%3Anetrunnereve%20OR%20author%3Afairydreaming)%20sort%3Aupdated-desc) / [dev stats](https://github.com/ggml-org/llama.cpp-dev) / [lib llama API](https://github.com/ggml-org/llama.cpp/issues/9289) / [llama-server REST API](https://github.com/ggml-org/llama.cpp/issues/9291)
## Quick start
A few options to get `llama.cpp` installed on your machine:
- Visit <https://llama.app> and follow the instructions
- Run with Docker - see our [Docker documentation](https://github.com/neurall/llama.cpp/blob/release/docs/docker.md)
- Download pre-built binaries from the [releases page](https://github.com/ggml-org/llama.cpp/releases)
- Build from source by cloning this repository - check out [our build guide](https://github.com/neurall/llama.cpp/blob/release/docs/build.md)
Once installed:
```
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
```
| | |
| --- | --- |
| [VLM session with `llama cli`](https://private-user-images.githubusercontent.com/1991296/629069102-88726b48-1713-48aa-a525-95a02e78afc4.png?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3OTAzNTQxMTUsIm5iZiI6MTc5MDM1MzgxNSwicGF0aCI6Ii8xOTkxMjk2LzYyOTA2OTEwMi04ODcyNmI0OC0xNzEzLTQ4YWEtYTUyNS05NWEwMmU3OGFmYzQucG5nP1gtQW16LUFsZ29yaXRobT1BV1M0LUhNQUMtU0hBMjU2JlgtQW16LUNyZWRlbnRpYWw9QUtJQVZDT0RZTFNBNTNQUUs0WkElMkYyMDI2MDkyNSUyRnVzLWVhc3QtMSUyRnMzJTJGYXdzNF9yZXF1ZXN0JlgtQW16LURhdGU9MjAyNjA5MjVUMTYzMDE1WiZYLUFtei1FeHBpcmVzPTMwMCZYLUFtei1TaWduYXR1cmU9NDkxZjEyNDVhY2E0ODQ4MWI1OWZhYWIxMmNiNGM4MWJhZjRlMjc3YmNjNzdhZThhZjkyMmVlMjk5OTIwZjYzNSZYLUF