HotPin reportedly released a small llama.cpp patch, initially named ExpertCache, that uses routing-guided memory locking or expert streaming to run large mixture-of-experts models with constrained consumer memory. The supplied results establish that 30B–120B MoE models can run on consumer-class systems under other configurations, including 24GB VRAM paired with 64GB system RAM, but they do not independently benchmark HotPin’s patch or substantiate lossless operation on roughly 24GB of RAM alone. Practical throughput, storage-I/O demands, and SSD-wear implications therefore remain unverified in the supplied material.
No intersection found in Scott’s wikis, and the radar does not already track this development or its actors. The case may fit Scott’s general local-inference interests, but without supporting hits there is no grounded basis for a stronger relevance claim.
queries asked of Scott's wikis
- local inference under consumer memory constraints
- MoE expert streaming and routing-guided caching
- lossless inference versus quantization tradeoffs
- disk-backed model inference and SSD wear
- llama.cpp performance patches and benchmark standards
- local model hardware economics
2026-08-05T20:33:51Z
The latest attachment again advances only the broader feasibility of MoE offloading, while HotPin’s specific patch remains without direct replication or meaningful upstream traction. Repeated adjacent evidence no longer justifies an active case; reopen only if independent benchmarks or llama.cpp adoption appear.
2026-08-05T18:32:55Z
No new material independently tests HotPin; the attached evidence again supports only the broader feasibility of MoE offloading. Repeated adjacent amplification no longer warrants frequent review, so wait for a direct replication or substantive upstream activity.
2026-08-05T16:35:30Z
The latest attachment still offers only adjacent evidence that MoE offloading is feasible, not an independent test of HotPin’s patch or its lossless-memory, speed, I/O, and SSD-wear claims. The case remains cold and uncorroborated pending direct replication.
2026-08-05T15:27:54Z
The attached evidence remains an adjacent MoE offload implementation, not an independent replication of HotPin’s patch or its memory, speed, I/O, and storage-wear claims. This is repetitive amplification of technical plausibility rather than new validation, so the case stays cold pending direct benchmarks.
2026-08-05T14:31:38Z
The added evidence remains an adjacent CPU-offload implementation rather than an independent test of HotPin, so it does not change the core claim’s evidentiary status. Direct replication of memory use, practical throughput, NVMe I/O, and storage wear is still required.
2026-08-05T13:32:15Z
The separate TensorSharp implementation makes RAM-resident MoE offload more technically plausible, but it does not benchmark HotPin or validate its lossless 24GB, throughput, I/O, or SSD-wear claims. Sparse, skeptical discussion adds no substantive corroboration, so this remains a low-temperature watch for direct replication.
2026-08-05T13:21:44Z
evidence attached: reddit.post.1vg71ci — A separate CPU-offload implementation with benchmark results provides useful independent evidence about practical RAM-resident MoE inference.
2026-07-26T05:21:20Z
The newly attached material still traces to the author’s original release and adds no independent benchmark or implementation report. The lossless memory, throughput, I/O, and storage-wear claims therefore remain uncorroborated.
2026-07-25T20:23:24Z
grounded: novel/none — No intersection found in Scott’s wikis, and the radar does not already track this development or its actors. The case may fit Scott’s general local-inference in
2026-07-25T20:22:44Z
origin walked (codex/luna, conf 0.99): anchor hn.story.49050356 -> echo.github.656083b29b by Ibrahim Khaled
2026-07-25T20:21:42Z
case created — HotPin presents a distinct, bounded implementation claim that is technically consequential but currently supported only by its author's lightly observed report.