John-194 has proposed a ggml-org/llama.cpp pull request adding an `--n-cpu-ffn` control that would place selected feed-forward-network components of dense models on the CPU, analogous to existing CPU-offload techniques for MoE models. The intended tradeoff is lower GPU VRAM usage in exchange for additional CPU work and data movement, potentially allowing larger quantized GGUF models to run on constrained local hardware. The supplied snippets establish that quantization and CPU/GPU partitioning can reduce VRAM pressure, especially for MoE models, but they do not independently benchmark this new dense-model option or establish its practical throughput.
The proposed control directly implements Scott’s hardware-aware local-inference position: accelerator placement and memory pressure should be explicit runtime policy, evaluated at the VRAM/throughput operating point. It could extend the models practical on his self-hosted GPU substrate, but the PR remains unbenchmarked, so it is currently an actionable evaluation target rather than an established improvement.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.evaluation-driven-developmentip:concept.operating-pointradar:concept.llama-cppradar:concept.local-inferenceradar:llama-cpp-hot-expert-gpu-cache
queries asked of Scott's wikis
- hybrid CPU GPU inference partitioning
- local inference VRAM versus throughput economics
- memory bandwidth bottlenecks in local LLM inference
- hardware-aware quantized model deployment
- dense FFN offload versus layer offload
- benchmark standards for local inference
2026-08-30T13:31:00Z
Repeated staleness now indicates the merged feature has not generated independent benchmarks, reproducible operating-point evidence, or visible adoption momentum. The implementation remains available for future rediscovery, but this episode has faded as an active developing story.
2026-08-28T13:29:19Z
New comments add only repetitive use-case and configuration discussion, with no independent VRAM/throughput benchmark or evidence of advantage over `-fit`. The merged feature remains testable but practically unvalidated.
2026-08-27T20:45:13Z
The refreshed discussion still offers no independent hardware benchmark, reproducible VRAM/throughput result, or evidence that manual FFN placement beats `-fit`. It is repetitive amplification of an available but practically unvalidated feature.
2026-08-27T12:30:09Z
The refreshed comments remain repetitive implementation discussion and add no independent measurement of VRAM savings, throughput, or advantage over `-fit`. The merged feature stays a concrete evaluation target, but its practical operating point remains unvalidated.
2026-08-27T11:31:17Z
Refreshed comments clarify the intended low-VRAM use case but continue to question whether manual FFN placement improves on `-fit`; they provide no independent measurements. The merged feature remains a testable optimization, not yet evidence of a better VRAM/throughput operating point.
2026-08-27T10:24:33Z
The reported upstream merge moves this from a proposal to an available, testable llama.cpp feature, raising its maturity despite the confirmation being carried through Reddit. Its practical value remains unsettled until hardware-specific tests compare VRAM savings and throughput against `-fit` and conventional layer offload.
2026-08-27T10:23:22Z
evidence attached: reddit.post.1vzp4c9 — shared external link with case evidence
2026-08-26T02:30:08Z
Minor Reddit engagement adds no validation or adoption signal; the proposed offload remains an unbenchmarked implementation awaiting an independent hardware test or upstream disposition. The hot local-inference backdrop does not change this case’s maturity.
2026-08-24T02:23:21Z
The staleness trigger adds no evidence: there is still no independent hardware benchmark, upstream merge, or reproducible VRAM/throughput result. The proposal remains a concrete but unvalidated evaluation target rather than a developing adoption signal.
2026-08-22T01:28:38Z
The case has only minor engagement drift and no independent benchmark, merge, or reproducible hardware result. Its meaning is unchanged: a concrete but still unvalidated local-inference optimization awaiting evidence of the VRAM/throughput tradeoff.
2026-08-20T01:25:07Z
The refreshed discussion remains implementation-level debate rather than independent validation; no reproducible VRAM/throughput result or merge changes the case’s meaning. Keep it as an evaluation target pending a hardware-specific benchmark or upstream adoption.
2026-08-19T12:32:24Z
The refreshed discussion adds implementation questions and author follow-up but no independent benchmark or reproducible VRAM/throughput result. The case remains a concrete evaluation target, while repetitive amplification warrants cooling until a merge or hardware-specific test appears.
2026-08-19T12:28:06Z
grounded: converges/medium — The proposed control directly implements Scott’s hardware-aware local-inference position: accelerator placement and memory pressure should be explicit runtime p
2026-08-19T12:25:25Z
case created — An open first-party implementation and reported Qwen3.8-27B results make this a concrete, testable local-inference episode.