2026-10-11 18:04 UTC

Independent benchmarks will determine whether HotPin's llama.cpp patches can run 30B–120B MoE models losslessly on roughly 24GB of consumer RAM at practical speeds without excessive storage wear.

state: expiredheat: lowuncertainty: highnovelscott: nonelocal-inference moe-models expert-streamingHotPinllama.cpp

What is this?

HotPin reportedly released a small llama.cpp patch, initially named ExpertCache, that uses routing-guided memory locking or expert streaming to run large mixture-of-experts models with constrained consumer memory. The supplied results establish that 30B–120B MoE models can run on consumer-class systems under other configurations, including 24GB VRAM paired with 64GB system RAM, but they do not independently benchmark HotPin’s patch or substantiate lossless operation on roughly 24GB of RAM alone. Practical throughput, storage-I/O demands, and SSD-wear implications therefore remain unverified in the supplied material.

Why it matters to Scott

No intersection found in Scott’s wikis, and the radar does not already track this development or its actors. The case may fit Scott’s general local-inference interests, but without supporting hits there is no grounded basis for a stronger relevance claim.
queries asked of Scott's wikis
  • local inference under consumer memory constraints
  • MoE expert streaming and routing-guided caching
  • lossless inference versus quantization tradeoffs
  • disk-backed model inference and SSD wear
  • llama.cpp performance patches and benchmark standards
  • local model hardware economics

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnAsk HN: HotPin – lossless 120B MoE inference on 24GB RAM (CPU, 50 loc)LozzKappa70
🟧 echo.github ⭐The earliest primary artifact is the author's initial GitHub release, originally named ExpertCache. Its commit states: “routing-guided mlockIbrahim Khaled——
🟠 redditMoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp
LocalLLaMA
fuzhongkai2339

Interpretation history

Decision trace