llama.cpp is a local-inference framework that supports splitting MoE model computation across CPU and GPU when VRAM is limited. The reported PR adds a CUDA-only heatmap that keeps frequently selected (“hot”) experts in GPU memory while computing colder experts on CPU, with its evidence title claiming a 33→56 tok/s improvement on an 8GB GPU. The supplied results support expert caching as a broader optimization—HOBBIT, a separate system built atop llama.cpp, reports up to 9.93× faster decoding—but do not independently benchmark this PR across models and quantizations, so its general speedup and regression profile remain unestablished here.
The PR converges with Scott’s hardware-aware local-inference position by turning expert placement and VRAM pressure into dynamic runtime policy; if independently validated across models and quantizations, it could directly inform experiments on his gamepc CUDA substrate. The radar tracks adjacent llama.cpp and MoE expert-streaming work, but not this specific hot-expert cache development, so this extends rather than repeats the tracked story.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.evaluation-driven-developmentip:concept.usable-mass-over-unusable-powerradar:concept.llama-cppradar:concept.moe-inferenceradar:concept.expert-streamingradar:person.llama-cppradar:hotpin-lossless-moe-streaming
queries asked of Scott's wikis
- MoE expert caching and CPU-GPU offload
- local inference on memory-constrained GPUs
- llama.cpp performance and optimization work
- quantization-dependent inference regressions
- benchmark methodology for local model runtimes
- consumer hardware economics for large MoE models
2026-08-07T22:28:23Z
After 48 hours, the case still has only the originating PR’s mixed results and trivial engagement drift, with no independent benchmark or implementation evidence. Let this episode expire and reopen only if cross-hardware, model, or quantization validation appears.
2026-08-05T22:24:16Z
The latest trigger is another empty re-observation, leaving the case entirely dependent on the originating PR’s mixed, model-specific results. Suppress engagement-driven review and revisit only when independent cross-hardware, model, or quantization benchmarks arrive.
2026-08-05T17:29:06Z
The latest trigger adds no substantive evidence and leaves the case dependent on the originating PR’s mixed, model-specific results. Further engagement-only reobservations should not prompt review; wait for independent cross-hardware, model, or quantization benchmarks.
2026-08-05T14:29:18Z
The new trigger is another empty re-observation, not independent validation. The case remains dormant and author-dependent; revisit only when cross-hardware, model, or quantization benchmarks test both gains and regressions.
2026-08-05T13:31:15Z
The new attachment is an empty re-observation, adding no independent benchmark, implementation result, or regression analysis. Repeated engagement triggers are exhausted; keep the case dormant until cross-hardware, model, or quantization evidence appears.
2026-08-05T12:23:30Z
No substantive evidence arrived; repeated engagement-only triggers have exhausted their informational value. Keep the case dormant until independent cross-hardware, model, or quantization benchmarks test both the reported gains and regressions.
2026-08-05T10:24:04Z
No independent benchmark or implementation evidence has arrived; this is continued amplification of the same mixed, author-reported results. Keep the case dormant until cross-hardware, model, or quantization testing materially changes the evidence base.
2026-08-05T09:27:15Z
No substantive evidence arrived; the trigger is another empty re-observation rather than independent validation. Keep the case dormant until cross-hardware, model, or quantization benchmarks clarify both gains and regressions.
2026-08-05T08:28:51Z
No substantive evidence arrived; this is another empty re-observation of the originating PR and its amplification. The case should remain dormant until independent cross-hardware, model, or quantization benchmarks test both gains and regressions.
2026-08-05T07:23:35Z
The attachment is another empty re-observation, leaving the case dependent on the originating PR’s mixed results rather than independent validation. Stop engagement-driven checks and revisit only for cross-hardware, model, or quantization benchmarks.
2026-08-05T06:26:41Z
The new attachment is only another re-observation of the same PR and Reddit discussion, so it adds no independent validation or regression evidence. Engagement-driven repetition is exhausted; hold the case until cross-hardware, model, or quantization benchmarks appear.
2026-08-05T05:23:44Z
The attached evidence is another re-observation of the originating PR and Reddit discussion, not independent validation. The mixed, model-dependent results remain unresolved, and further engagement-only triggers should be ignored until cross-hardware, model, or quantization benchmarks appear.
2026-08-05T04:22:59Z
The trigger adds no independent benchmark or implementation evidence; the case remains an author-reported, model-dependent optimization with unresolved regressions. Repeated amplification is no longer informative, so wait for cross-hardware, model, or quantization validation.
2026-08-05T03:27:20Z
The newly attached evidence still traces to the originating PR and Reddit amplification, adding no independent benchmark or implementation result. Repeated re-observation is not changing the case; wait for cross-hardware, model, and quantization validation.
2026-08-05T02:29:36Z
The attachment still adds no independent benchmark or implementation evidence beyond the originating PR and its amplification. The case remains unresolved but does not warrant repeated engagement-driven review; wait for cross-hardware, model, and quantization results.
2026-08-05T01:22:01Z
The new attachment adds no independent evidence beyond the originating PR and its Reddit amplification. Repeated engagement updates do not alter the mixed, model-dependent result; revisit only when cross-hardware, model, or quantization benchmarks appear.
2026-08-05T00:26:42Z
The attachment adds no independent benchmark or implementation evidence and does not change the case beyond the originating PR’s mixed results. Repeated amplification no longer merits hourly review; wait for cross-hardware, model, or quantization validation.
2026-08-04T23:28:44Z
The latest trigger adds no independent benchmark or implementation evidence, so the case remains anchored to the author’s mixed, model-dependent results. Repeated engagement-only updates are not changing its meaning; wait for cross-hardware, model, or quantization validation.
2026-08-04T22:27:40Z
The latest attachment adds no independent benchmark or implementation evidence; discussion remains amplification of the originating PR and repeats known concerns about output variation and extreme quantization. The optimization is still promising but unvalidated across hardware, models, and quantizations.
2026-08-04T21:23:52Z
The newly attached material still resolves to the originating PR and its Reddit amplification, not an independent benchmark. The case remains promising but unchanged: gains are model-dependent, with reported slowdowns and output-variation concerns still requiring validation across hardware and quantizations.
2026-08-04T20:24:39Z
The attached evidence still traces back to the same PR and Reddit amplification, with no independent benchmark or implementation result. Model-dependent slowdowns and output-variation concerns keep the optimization promising but unvalidated across hardware, models, and quantizations.
2026-08-04T19:26:59Z
The new activity is only marginal amplification of the original report; no independent benchmark, implementation result, or regression analysis has arrived. The claimed gains remain model-dependent and unvalidated across hardware and quantizations.
2026-08-04T18:25:54Z
grounded: converges/medium — The PR converges with Scott’s hardware-aware local-inference position by turning expert placement and VRAM pressure into dynamic runtime policy; if independentl
2026-08-04T18:23:33Z
origin walked (codex/luna, conf 0.99): anchor reddit.post.1vfhns3 -> echo.github.faa30ff3e3 by miltos22
2026-08-04T18:22:08Z
case created — The reported 1.7–2.1x gains and model-dependent negative results make this a concrete, independently resolvable inference optimization.