2026-10-11 17:09 UTC

Independent benchmarks will determine whether llama.cpp’s hot-expert GPU cache materially accelerates CPU-offloaded MoE inference on memory-constrained GPUs without regressions across models and quantizations.

state: expiredheat: lowuncertainty: highconvergesscott: mediumlocal-inference moe-inference llama-cppllama.cpp

What is this?

llama.cpp is a local-inference framework that supports splitting MoE model computation across CPU and GPU when VRAM is limited. The reported PR adds a CUDA-only heatmap that keeps frequently selected (“hot”) experts in GPU memory while computing colder experts on CPU, with its evidence title claiming a 33→56 tok/s improvement on an 8GB GPU. The supplied results support expert caching as a broader optimization—HOBBIT, a separate system built atop llama.cpp, reports up to 9.93× faster decoding—but do not independently benchmark this PR across models and quantizations, so its general speedup and regression profile remain unestablished here.

Why it matters to Scott

The PR converges with Scott’s hardware-aware local-inference position by turning expert placement and VRAM pressure into dynamic runtime policy; if independently validated across models and quantizations, it could directly inform experiments on his gamepc CUDA substrate. The radar tracks adjacent llama.cpp and MoE expert-streaming work, but not this specific hot-expert cache development, so this extends rather than repeats the tracked story.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.evaluation-driven-developmentip:concept.usable-mass-over-unusable-powerradar:concept.llama-cppradar:concept.moe-inferenceradar:concept.expert-streamingradar:person.llama-cppradar:hotpin-lossless-moe-streaming
queries asked of Scott's wikis
  • MoE expert caching and CPU-GPU offload
  • local inference on memory-constrained GPUs
  • llama.cpp performance and optimization work
  • quantization-dependent inference regressions
  • benchmark methodology for local model runtimes
  • consumer hardware economics for large MoE models

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditA llama.cpp PR caches “hot” MoE experts on the GPU — 33 → 56 tok/s reported with 8GB VRAM
LocalLLaMA
BTA_Labs26253
🟧 echo.github ⭐The PR says it is a CUDA-only feature that tracks an expert-usage heatmap, caches the hottest experts on GPU, and computes cold experts on Cmiltos22——

Interpretation history

Decision trace