llama.cpp contributor ngxson proposed a ggml-org/llama.cpp model-loader change, described in the supplied evidence titles as lazy mmap row/tensor loading for large tensors. The claimed benefit is that sparse Qwen-family models need load only tensors or rows actually used, potentially allowing local inference with substantially less RAM or VRAM. The search snippets establish that model storage, KV cache, and memory bandwidth constrain llama.cpp inference, but they do not independently quantify this patch’s savings or confirm its merge status and production performance.
The radar already tracks essentially the same sparse-MoE memory-reduction thesis in `radar:hotpin-lossless-moe-streaming`, with adjacent llama.cpp expert-streaming work also under observation. This implementation could affect Scott’s hardware-aware local-inference stack and gamepc/Ollama model capacity, but the supplied evidence does not confirm merge status, savings, or usable throughput.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaradar:hotpin-lossless-moe-streamingradar:concept.expert-streamingradar:concept.llama-cpp
queries asked of Scott's wikis
- lazy loading for sparse or mixture-of-experts models
- local inference memory economics
- mmap and demand paging in model runtimes
- consumer hardware as an AI capability frontier
- llama.cpp runtime optimizations
- RAM versus VRAM tradeoffs for local models
2026-08-31T17:34:38Z
A second staleness window passed without a memory benchmark, merge, residency analysis, or independent reproduction. The proposal remains plausible but has not developed into a distinct episode beyond the broader sparse-model streaming thesis, so passive sensors can reopen it on a material implementation result.
2026-08-29T17:27:54Z
No new technical evidence arrived within the staleness window; the implementation remains an early, credible proposal with one performance anecdote but no measured memory savings, merge, or reproducible validation. Keep watching, but the case has not advanced beyond its overlap with the broader sparse-model streaming thesis.
2026-08-27T16:34:46Z
A hands-on user now reports a warmed-cache throughput increase from roughly 11 to 13 t/s, suggesting the lazy-loading path can run without an obvious speed penalty in at least one constrained setup. This advances the case into early testing, but the report does not measure working-set reduction, residency behavior, or broad model support, so the core memory claim remains uncorroborated.
2026-08-27T15:43:57Z
The new activity is engagement-only and adds no evidence on merge status, memory savings, residency behavior, or throughput. The concrete proposal remains worth monitoring, but it has not advanced beyond an unvalidated implementation overlapping an existing sparse-model memory thesis.
2026-08-27T15:29:57Z
grounded: known/medium — The radar already tracks essentially the same sparse-MoE memory-reduction thesis in `radar:hotpin-lossless-moe-streaming`, with adjacent llama.cpp expert-stream
2026-08-27T15:27:54Z
origin walked (codex/luna, conf 0.99): anchor reddit.post.1vzw3jr -> echo.github.446671c590 by Xuan Son Nguyen
2026-08-27T15:26:47Z
case created — The linked first-party pull request is a concrete implementation targeting a distinct sparse-model memory bottleneck not covered by the existing dense FFN-offload case.