2026-10-11 17:10 UTC

Independent benchmarks will determine whether llama.cpp’s proposed dense-model CPU-FFN offload materially reduces VRAM requirements while preserving practical throughput for large quantized models.

state: expiredheat: lowuncertainty: highconvergesscott: mediumllama-cpp local-inference gpu-offloadJohn-194ggml-org

What is this?

John-194 has proposed a ggml-org/llama.cpp pull request adding an `--n-cpu-ffn` control that would place selected feed-forward-network components of dense models on the CPU, analogous to existing CPU-offload techniques for MoE models. The intended tradeoff is lower GPU VRAM usage in exchange for additional CPU work and data movement, potentially allowing larger quantized GGUF models to run on constrained local hardware. The supplied snippets establish that quantization and CPU/GPU partitioning can reduce VRAM pressure, especially for MoE models, but they do not independently benchmark this new dense-model option or establish its practical throughput.

Why it matters to Scott

The proposed control directly implements Scott’s hardware-aware local-inference position: accelerator placement and memory pressure should be explicit runtime policy, evaluated at the VRAM/throughput operating point. It could extend the models practical on his self-hosted GPU substrate, but the PR remains unbenchmarked, so it is currently an actionable evaluation target rather than an established improvement.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.evaluation-driven-developmentip:concept.operating-pointradar:concept.llama-cppradar:concept.local-inferenceradar:llama-cpp-hot-expert-gpu-cache
queries asked of Scott's wikis
  • hybrid CPU GPU inference partitioning
  • local inference VRAM versus throughput economics
  • memory bandwidth bottlenecks in local LLM inference
  • hardware-aware quantized model deployment
  • dense FFN offload versus layer offload
  • benchmark standards for local inference

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit[Open PR] llama : add --n-cpu-ffn option by John-194 · Pull Request #26622 · ggml-org/llama.cpp
LocalLLaMA
pmttyji3412
🟧 echo.github ⭐The pull request adds CPU-FFN offload controls for dense models, analogous to llama.cpp’s existing MoE CPU-offload options.John-194——
🟠 redditllama : add --n-cpu-ffn option by John-194 · Pull Request #26622 · ggml-org/llama.cpp
LocalLLaMA
jacek20239027

Interpretation history

Decision trace