A llama.cpp fork attributed in the case to timadinorth selectively keeps frequently activated mixture-of-experts (MoE) experts in GPU VRAM while leaving others in system RAM. The author claims this raised generation throughput from 20 to 30 tokens per second on partially offloaded coding, refactoring, and code-review workloads. The supplied web snippets support the general mechanism—CPU/GPU expert placement trades scarce VRAM for lower transfer latency—but do not independently verify this fork, the workload methodology, or the claimed 50% gain.
The radar already tracks this exact development in `radar:llama-cpp-hot-expert-gpu-cache`. It bears directly on Scott’s hardware-aware placement policy and memory-constrained `gamepc` inference stack, but the claimed gain remains unverified and adds nothing beyond the existing open case.
dev:concept.hardware-aware-local-inferencedev:project.gamepcradar:llama-cpp-hot-expert-gpu-cache
queries asked of Scott's wikis
- hot-expert caching for MoE inference
- adaptive CPU GPU model offloading
- local inference under VRAM constraints
- llama.cpp performance optimization projects
- workload-specific expert routing stability
- local coding-agent inference economics
2026-08-29T21:30:55Z
Adjacent implementations and the active llama.cpp RFC now strengthen the general hot-expert caching mechanism, but still do not reproduce this fork’s 20-to-30 t/s result. This duplicate episode is superseded by the existing canonical radar case, where any upstream adoption or independent benchmark should be tracked.
2026-08-29T19:25:08Z
evidence attached: hn.story.49492409 — The artifact offers independent evidence for selective MoE expert caching across VRAM, RAM, and NVMe to make oversized local models usable.
2026-08-29T19:25:07Z
evidence attached: reddit.post.1w1uu6d — The roundup documents active llama.cpp work on hot-expert caching and related CPU, RAM, disk, and hybrid-inference optimizations.
2026-08-29T18:31:50Z
The refreshed comments remain repetitive support for the general hot-expert caching mechanism, without independently reproducing this fork’s 20-to-30 t/s result or indicating an upstream llama.cpp decision. The case remains an unverified, workload-specific claim already covered by the existing radar episode.
2026-08-29T15:34:37Z
The refreshed discussion still offers only anecdotal results from related forks and analogous implementations, not an independent reproduction of this fork’s 20-to-30 t/s result or an upstream llama.cpp decision. The case remains a plausible but workload-specific claim already covered by the existing radar episode.
2026-08-29T12:29:12Z
The refreshed discussion still adds only analogous implementations and mechanism-level context, not an independent reproduction or upstream adoption. The 20-to-30 t/s result remains a workload-specific single-author claim already covered by the existing radar case.
2026-08-29T09:25:54Z
The refreshed comments remain repetitive mechanism-level context and provide no independent benchmark, reproducible methodology, or upstream implementation signal. The workload-specific 50% gain therefore remains a single-author claim already represented by the existing radar case.
2026-08-29T04:30:10Z
The refreshed discussion places the fork alongside an existing llama.cpp expert-cache RFC and analogous approaches, strengthening the mechanism’s plausibility but not the claimed 50% gain. No independent benchmark, reproducible methodology, or upstream adoption changes the case’s meaning.
2026-08-29T02:28:52Z
No new evidence or independent reproduction has appeared; the quantified gain remains a single-author workload-specific claim already covered by an existing radar case.
2026-08-29T02:26:27Z
grounded: known/medium — The radar already tracks this exact development in `radar:llama-cpp-hot-expert-gpu-cache`. It bears directly on Scott’s hardware-aware placement policy and memo
2026-08-29T02:24:35Z
case created — The linked implementation and quantified throughput claim define a concrete, distinct local-inference optimization episode.