2026-10-11 17:10 UTC

The fork author claims selectively placing frequently used MoE experts in VRAM raises llama.cpp generation throughput from 20 to 30 tokens per second on partially offloaded coding workloads, potentially improving local inference on memory-constrained GPUs.

state: resolvedheat: lowuncertainty: highknownscott: mediumlocal-inference moe llama-cpptimadinorthllama.cpp

What is this?

A llama.cpp fork attributed in the case to timadinorth selectively keeps frequently activated mixture-of-experts (MoE) experts in GPU VRAM while leaving others in system RAM. The author claims this raised generation throughput from 20 to 30 tokens per second on partially offloaded coding, refactoring, and code-review workloads. The supplied web snippets support the general mechanism—CPU/GPU expert placement trades scarce VRAM for lower transfer latency—but do not independently verify this fork, the workload methodology, or the claimed 50% gain.

Why it matters to Scott

The radar already tracks this exact development in `radar:llama-cpp-hot-expert-gpu-cache`. It bears directly on Scott’s hardware-aware placement policy and memory-constrained `gamepc` inference stack, but the claimed gain remains unverified and adds nothing beyond the existing open case.
dev:concept.hardware-aware-local-inferencedev:project.gamepcradar:llama-cpp-hot-expert-gpu-cache
queries asked of Scott's wikis
  • hot-expert caching for MoE inference
  • adaptive CPU GPU model offloading
  • local inference under VRAM constraints
  • llama.cpp performance optimization projects
  • workload-specific expert routing stability
  • local coding-agent inference economics

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit50% tg increase with offloading "hot" experts to VRAM
LocalLLaMA
nbvehrfr3420
🟧 echo.github ⭐A llama.cpp change selectively offloads experts observed to remain frequently used across coding, refactoring, and code-review workloads.timadinorth——
🟠 redditllama.cpp Open PRs list - CPU/RAM/Disk/Hybrid Related - Better for CPU-only & Hybrid inference
LocalLLaMA
pmttyji15939
🟧 hnShow HN: Moe-Direct – MoE Models far larger than your RAM, on a consumer desktoptmxkzm1925-max10

Interpretation history

Decision trace