2026-10-11 16:34 UTC

Fork author neuralll claims his released llama.cpp fork's VRAM-filling hot-expert cache (built on csantiago78's PR #27861) roughly doubles decode throughput for GLM-5.3-Flash and MiMo MoE models far larger than total VRAM on two RTX 3090s with unchanged perplexity, and projects further gains per added GPU โ€” if independent multi-GPU users reproduce it, hot-expert caching becomes a practical standard path for memory-constrained local MoE inference.

state: acceleratingheat: mediumuncertainty: mediumconvergesscott: highlocal-inference llama-cpp moe-offloading multi-gpu expert-cachingneuralllcsantiago78

What is this?

neuralll published a llama.cpp fork extending csantiago78's open PR #27861 โ€” a GPU-resident LRU cache for MoE expert weights offloaded to system RAM โ€” to fill all free VRAM across multiple GPUs with hot experts; on two RTX 3090s (48 GB combined, per the hardware tables in the snippets) he reports roughly 2x decode for GLM-5.3-Flash and MiMo quants far larger than that VRAM at unchanged perplexity, with per-GPU gains projected but untested. No snippet reproduces neuralll's own numbers, but the underlying cache is now independently measured on the same hardware class: a writeup showing ~17โ†’25-29 tok/s on 2x RTX 3090 with a 157 GiB Qwen3.8-Flash-Next quant, a club-3090 discussion running 118B/284B models on dual 3090s, and a dual-AMD-Vulkan report going ~8โ†’19 tok/s at ~93% hit rate before spilling. GLM-5.3-Flash itself (Z.ai's open-weight 320B-total/18B-active MoE, ~190 GB at practical 4-bit) is the anchor use case, and rival >VRAM paths are surfacing in the same window: LayerStoRm's expert-streaming engine claims GLM-5.3-Flash 186 GiB on 96 GB VRAM at ~24.5 tok/s, and a separate single-GPU expert hot cache (PR #26563 lineage) beats layer-level --n-cpu-moe tuning on a simulated 5090. So the mechanism's premise โ€” routing locality making hot-expert caching pay โ€” now has independent corroboration across vendors, while neuralll's specific fork numbers remain first-party only.

Why it matters to Scott

The fork's core premise โ€” routing locality making hot-expert caching pay โ€” now has independent multi-GPU corroboration on Scott's exact hardware class (dual-3090 ~17โ†’25-29 tok/s with a 157 GiB Qwen quant, dual-AMD Vulkan ~8โ†’19 tok/s at ~93% hit rate), converging with his hardware-aware-local-inference position that expert placement and memory pressure are explicit runtime policy. neuralll's own numbers remain first-party-only, but with rival >VRAM paths (LayerStoRm expert-streaming, single-GPU hot-cache lineage) surfacing in the same window, the case for benching the fork on gamepc's dual 3090s and publishing dated receipts against radar:llama-cpp-hot-expert-gpu-cache materially strengthened.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaradar:llama-cpp-hot-expert-gpu-cacheradar:afm3-prompt-conditioned-pruningradar:adaptive-kv-cache-streamingradar:airllm-low-vram-model-streaming
queries asked of Scott's wikis
  • hardware-aware local inference expert placement memory pressure
  • MoE expert offloading host RAM bandwidth decode bottleneck
  • llama.cpp multi-GPU 3090 benchmark tok/s perplexity methodology
  • expert routing temporal locality cache hit rate
  • VRAM-bound MoE serving economics consumer GPU rigs
  • fork versus upstream patch merge tradeoffs

Measured heat

now 0 pts/hpeak 111 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 385h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-25 16:33 (minted)โญ origin echo-reconstructedREADME titled 'fork with multi gpu acceleration even for models bigger than total gpu mem': fills all free VRAM with a hot-expert cache of t
neuralll on github (echo) ยท attributed from hn.story.49845314 ยท published time unknown
โ€”
09-25 14:40first on hacker news ยท published ยท lag ?Show HN: Llama.cpp fork with 2-4x multiGPU speed for MoE models bigger than VRAM
neuralll
โ€”
10-06 12:29first on r/LocalLLaMA ยท published ยท lag ?Qwen3-Next-80B on a 3090 with 16GB RAM: ~3x faster decode than stock llama.cpp by not waiting for every expert (patch + paper)
Zestyclose_Reality15
โ€”
09-25 14:40amplified on hacker newshn.story.49845314
neuralll
peak 2 ยท 3 comments ยท 1% of case engagement
09-29 14:46amplified on hacker newshn.story.49894258
gslaller
peak 2 ยท 0 comments ยท 1% of case engagement
10-06 12:29amplified on r/LocalLLaMAreddit.post.1wz1bxs
Zestyclose_Reality15
peak 9 ยท 8 comments ยท 3% of case engagement
10-07 18:15amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1x03xkc
jacek2023
peak 445 ยท 139 comments ยท 95% of case engagement
09-25 15:21our radar first saw it ยท lag ?discovery anchor: hn.story.49845314โ€”
pace: p86 vs 1032 stories at the 336h mark (now 385h old) โ€” ahead of cloudflare-pingora-ketama-memory-reduction (1.0x), behind cactus-needle3-on-device-automation (1.0x)

Evidence (5) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnShow HN: Llama.cpp fork with 2-4x multiGPU speed for MoE models bigger than VRAM
Retrieved article excerpt

Open article ยท Retrieved 2026-09-25T16:30:16.172096+00:00

# llama.cpp: fork with multi gpu acceleration even for models bigger than total gpu mem

> **Big thanks to [@csantiago78](https://github.com/csantiago78)**: the expert cache
> here builds on their implementation in llama.cpp PR
> [#27861](https://github.com/ggml-org/llama.cpp/pull/27861) ("GPU-resident LRU cache
> for host-offloaded MoE expert weights"), the first to get a working hot-expert
> cache into llama.cpp. This fork takes it further: filling all free VRAM, smarter
> eviction, CPU/GPU overlap and fused kernels.

For MoE models much larger than VRAM: every expert stays in system RAM, and all
VRAM left after the KV cache becomes a live cache of the experts actually being
used. The GPUs compute cached experts while the CPU computes the rest, in parallel.

**Results, 2x RTX 3090 (48 GB) + 125 GB RAM, single stream, temp 0, `-c 1024`:**

| model | size | stock t/s | fork t/s | gain | PPL (stock -> fork) |
| --- | --- | --- | --- | --- | --- |
| GLM-5.3-Flash 3.0-bit, original GGUF | 117 GB | 12.3 | ~25 | 2.0x | 3.5534 -> 3.5534 |
| GLM-5.3-Flash 3.0-bit, [Q4\_K attention GGUF](https://huggingface.co/neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF) | 106 GB | 12.3 | **27.72** | 2.3x | 3.5534 -> 3.5871 (+0.95%) |
| MiMo-2.6-Flash-RL IQ3\_XXS | 132 GB | 4.7 | **9.9** | 2.1x | unchanged (no requant) |
| Qwen3.8-Flash-Next UD-IQ4\_XS | 88 GB | 29.7 | **32.1** | 1.08x | unchanged (no requant) |

GLM prompt: "generate smallest html tetris game."; MiMo/Qwen prompt: "write smallest
html tetris game" (both temp 0). PPL: wikitext-2, 40 x 512-token chunks (GLM only).
MiMo needed `-fitt 8000` on both stock and fork to avoid autofit OOM-ing on this
arch/quant combo; the others loaded fine with default fit.

Qwen's gain is small because stock's autofit already placed most of its experts on
GPU by default here โ€” little room left for the cache to improve on. The big wins
(GLM, MiMo) are on models where default placement leaves most expert work on the CPU.

Models:

- GLM original GGUF (tested): [pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF](https://huggingface.co/pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF),
  the 3.0-bit file. It works as-is with this fork, no conversion needed, and it's the
  quality reference (unchanged perplexity). Most of the speedup comes from the fork,
  not the requantization.
- GLM Q4\_K attention variant (same experts, non-expert Q8\_0 weights requantized to Q4\_K):
  [neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF](https://huggingface.co/neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF).

**Same VRAM, different use.** Stock llama.cpp and this fork get the same 48 GB; what
differs is what it holds. Each token uses only 8 of the 288 experts in each layer.

- Stock places experts statically, whole layers at a time: ~36 GB fits all 288
  experts of ~14 of the 42 MoE layers. Most of that VRAM holds experts the current
  token doesn't touch, so only ~33% of each token's expert work runs on GPU and the
  CPU does ~67%, one after the other.
- This fork fills the same VRAM with the ~100 most-used experts of every layer.
  Usage is skewed, so those cover ~85% of what tokens actually pick: ~85% of expert
  work runs on GPU and the CPU does ~15%, at the same time as the GPUs.

|  | expert work on GPU | expert work on CPU | decode t/s |
| --- | --- | --- | --- |
| stock (static whole layers) | ~33% | ~67% | 12.3 |
| this fork (cache of hot experts) | ~85% | ~15%, in parallel | 26-28 |

So stock can't reach 2x on the same hardware: without an expert cache, extra VRAM
mostly holds experts that aren't being used.

Run (with the faster Q4\_K attention GGUF from [neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF](https://huggingface.co/neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF)):

```
llama-server -m GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf \
    -np 1 -c 1024 -t 6 --cpu-moe -nr --moe-expert-cache -1
```

- `--cpu-moe` keeps all experts in RAM; `--moe-expert-cache -1` sizes the cache
  per GPU from the VRAM free after KV and compute buffers. Set `-c` explicitly:
  without it autofit grows the context and takes the VRAM the cache needs.
- `-t 6` suits an 8-core CPU (leave cores to drive the GPUs).
- Cache hit rate is ~85% after warm-up. `LLAMA_MOE_CACHE_STATS=1` logs it.
- Tuning: `LLAMA_MOE_CACHE_POLICY` (`add` default, `halve`, `window`),
  `LLAMA_MOE_CACHE_MARGIN_MB` (VRAM left free, default 1024),
  `LLAMA_MOE_CACHE_SWAP_FRAC` (share of token time for uploads, default 0.25).
- `GGML_SCHED_PROF=1` prints where each token's host time goes.

**More GPUs (estimate, only 2 tested).** Nothing assumes two GPUs: each GPU gets
its own cache, sized from its free VRAM. Each extra GPU adds cache room, so more
of every token's experts are hits and less work falls to the CPU. For this model,
per token today: ~20 ms GPU work on non-expert layers, ~10 ms CPU on missed experts.

| GPUs (24 GB each) | cache room | slots/layer (of 288) | hit rate | CPU miss time | decode t/s |
| --- | --- | --- | --- | --- | --- |
| 2 (measured) | ~34 GB | ~100 | 85-88% | ~10 ms | 26-28 |
| 3 | ~58 GB | ~170 | ~95% | ~3-4 ms | ~33-35 |
| 4 | ~82 GB | ~240 | ~99% | ~1 ms | ~38-42 |
| 5+ | whole model | 288 | 100% | 0 | ~40-45 (plateau) |

The plateau is the ~20 ms GPU part: with the default layer split each layer runs on
one GPU at a time, so extra GPUs add cache room, not speed on that part. System RAM
must still hold all experts. Cards in x4 PCIe slots upload experts slower, so the
cache warms up slower. Reports from 3+ GPU setups are welcome.

What's in it: the GPU expert cache from PR [#27861](https://github.com/ggml-org/llama.cpp/pull/27861)
(csantiago78), extended with VRAM-filling auto-sizing, usage-driven eviction that
only swaps when the upload pays back, CPU/GPU overlap per layer, scheduler barrier
fixes and fused gate kernels. GLM-5.3-Flash support comes from PRs
[#27773](https://github.com/ggml-org/llama.cpp/pull/27773) and
[#27917](https://github.com/ggml-org/llama.cpp/pull/27917) (timkhronos); stock
llama.cpp can't load GLM-5.3-Flash yet.

**Prefill warm start**, implemented independently for this fork: the cache observes
which experts the prompt itself selects during prefill and preloads them before the
first generated token, instead of starting cold and only learning from decode.
[@sdroege](https://github.com/sdroege) explored the same idea independently
in the [PR #27861 discussion](https://github.com/ggml-org/llama.cpp/pull/27861)
with their own patch; worth checking out too.

**Other notable tweaks:**

- **Swap budget from measured cost, not a guess**: each step measures real upload
  time (ms/expert) and real token time, then computes how many swaps fit in
  `LLAMA_MOE_CACHE_SWAP_FRAC` of a token (default 25%) โ€” instead of a fixed
  swaps-per-step constant.
- **Pay-back filter on every eviction**: a swap only happens if the candidate's
  measured usage beats the cached victim's by more than what the upload itself
  costs in CPU-equivalent time, so churn can't cost more than it saves.
- **Usage tracking with decay** (`LLAMA_MOE_CACHE_POLICY`: `add` default, `halve`,
  `window`): recent use counts more than old use, so the cache follows shifts in
  which experts are hot instead of freezing on early-token bias.
- **CPU/GPU overlap inside a layer**: the GPU cache chain is queued and its inputs
  copied *before* the CPU miss chain runs, so both compute at the same time instead
  of the scheduler serializing them.
- **Scheduler fixes upstream benefits from too**: no host barrier between two GPU
  splits when neither reads host memory, and a new split is inserted exactly when
  another GPU's result is needed mid-split โ€” both apply to any multi-GPU llama.cpp
  workload, not just this cache.
- **Fused CUDA kernel** for the hyper-connection/KDA gate chain
  (`MUL -> ADD|SCALE -> SIGMOID -> SCALE`, one kernel instead of four), and GLM5-Next's
  KDA Q/K norm collapsed from `rms_norm`+`scale` into one `l2_norm` op.
- **Repacked CPU experts stay cacheable**: uploads read raw bytes straight from the
  GGUF file (recorded per-tensor file offsets) when host memory holds a
  repack-transformed layout instead of the on-disk one.
- `GGML_SCHED_PROF=1` and `LLAMA_MOE_CACHE_STATS=1` for live profiling: barrier vs.
  copy wait time, fill %, in-flight uploads, queue depth, hit rate.

**About the author of this fork**: I'm actively looking for an AI engineering/research
role and open to relocating out of Eastern Europe. If this work is useful to you or
your team, reach out: [linkedin.com/in/neuralll](https://www.linkedin.com/in/neuralll/)

---

[llama](https://raw.githubusercontent.com/ggml-org/llama.brand/refs/heads/master/cover/llama-cpp/cover-llama-cpp-dark.svg)

**LLM inference in C/C++**

[License: MIT](https://opensource.org/licenses/MIT)
[Release](https://github.com/ggml-org/llama.cpp/releases?q=tag:v0)
[Nightly](https://github.com/ggml-org/llama.cpp/releases?q=b)
[Server](https://github.com/ggml-org/llama.cpp/actions/workflows/server.yml)
[Docker](https://github.com/ggml-org/llama.cpp/actions/workflows/docker.yml)
[Winget](https://github.com/ggml-org/llama.cpp/actions/workflows/winget.yml)

[ggml](https://github.com/ggml-org/ggml) / [ops](https://github.com/ggml-org/llama.cpp/blob/master/docs/ops.md) / [maintainer PRs](https://github.com/ggml-org/llama.cpp/issues?q=is%3Apr%20is%3Aopen%20draft%3AFalse%20(author%3Argerganov%20OR%20author%3AKitaitiMakoto%20OR%20author%3Adanbev%20OR%20author%3Aaldehir%20OR%20author%3Amax-krasnyansky%20OR%20author%3ACISC%20OR%20author%3Aggerganov%20OR%20author%3Aam17an%20OR%20author%3Ajhen0409%20OR%20author%3Abartowski1182%20OR%20author%3Anikwen%20OR%20author%3Ahipudding%20OR%20author%3Aravi9%20OR%20author%3AServeurpersoCom%20OR%20author%3Apwilkin%20OR%20author%3Areeselevine%20OR%20author%3Angxson%20OR%20author%3Ajeffbolznv%20OR%20author%3Amarty1885%20OR%20author%3A0cc4m%20OR%20author%3ATitaniumtown%20OR%20author%3Aangt%20OR%20author%3AIMbackK%20OR%20author%3Aarthw%20OR%20author%3AJohannesGaessler%20OR%20author%3AORippler%20OR%20author%3Aruixiang63%20OR%20author%3Axctan%20OR%20author%3Aallozaur%20OR%20author%3Ayomaytk%20OR%20author%3Aaendk%20OR%20author%3Awine99%20OR%20author%3Agaugarg-nv%20OR%20author%3Ataronaeo%20OR%20author%3Aforforever73%20OR%20author%3Alhez%20OR%20author%3Anetrunnereve%20OR%20author%3Afairydreaming)%20sort%3Aupdated-desc) / [dev stats](https://github.com/ggml-org/llama.cpp-dev) / [lib llama API](https://github.com/ggml-org/llama.cpp/issues/9289) / [llama-server REST API](https://github.com/ggml-org/llama.cpp/issues/9291)

## Quick start

A few options to get `llama.cpp` installed on your machine:

- Visit <https://llama.app> and follow the instructions
- Run with Docker - see our [Docker documentation](https://github.com/neurall/llama.cpp/blob/release/docs/docker.md)
- Download pre-built binaries from the [releases page](https://github.com/ggml-org/llama.cpp/releases)
- Build from source by cloning this repository - check out [our build guide](https://github.com/neurall/llama.cpp/blob/release/docs/build.md)

Once installed:

```
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
```

|  |  |
| --- | --- |
| [VLM session with `llama cli`](https://private-user-images.githubusercontent.com/1991296/629069102-88726b48-1713-48aa-a525-95a02e78afc4.png?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3OTAzNTQxMTUsIm5iZiI6MTc5MDM1MzgxNSwicGF0aCI6Ii8xOTkxMjk2LzYyOTA2OTEwMi04ODcyNmI0OC0xNzEzLTQ4YWEtYTUyNS05NWEwMmU3OGFmYzQucG5nP1gtQW16LUFsZ29yaXRobT1BV1M0LUhNQUMtU0hBMjU2JlgtQW16LUNyZWRlbnRpYWw9QUtJQVZDT0RZTFNBNTNQUUs0WkElMkYyMDI2MDkyNSUyRnVzLWVhc3QtMSUyRnMzJTJGYXdzNF9yZXF1ZXN0JlgtQW16LURhdGU9MjAyNjA5MjVUMTYzMDE1WiZYLUFtei1FeHBpcmVzPTMwMCZYLUFtei1TaWduYXR1cmU9NDkxZjEyNDVhY2E0ODQ4MWI1OWZhYWIxMmNiNGM4MWJhZjRlMjc3YmNjNzdhZThhZjkyMmVlMjk5OTIwZjYzNSZYLUF
neuralll23
๐ŸŸง echo.github โญREADME titled 'fork with multi gpu acceleration even for models bigger than total gpu mem': fills all free VRAM with a hot-expert cache of tneuralllโ€”โ€”
๐ŸŸง hnShow HN: Moe Routing Atlas โ€“ which experts fire in Qwen3.5/3.6 35B-A3Bgslaller20
๐ŸŸ  redditQwen3-Next-80B on a 3090 with 16GB RAM: ~3x faster decode than stock llama.cpp by not waiting for every expert (patch + paper)
LocalLLaMA
Zestyclose_Reality1598
๐ŸŸ  redditllama : add a GPU cache for MoE experts kept in host memory by am17an ยท Pull Request #29887 ยท ggml-org/llama.cpp
LocalLLaMA
jacek2023445139

Interpretation history

Decision trace