2026-10-11 17:13 UTC

The MoE Offload Bench maintainer claims the released implementation can offload sparse-model experts on a two-core Celeron with 2.7GB of RAM, potentially extending local MoE inference to extremely constrained commodity systems.

state: expiredheat: lowuncertainty: highknownscott: mediumlocal-inference inference-economics open-modelsdhishwasher

What is this?

The case reports a released MoE expert-offloading implementation whose maintainer, identified as dhishwasher, claims it can run on a two-core Celeron with 2.7GB of RAM. The returned snippets substantiate the underlying approach: llama.cpp can place MoE expert tensors in CPU RAM, while sparse models activate only a subset of experts per token; related releases add Rust bindings and packaged model integrations. However, the snippets do not independently identify the claimed benchmark repository or verify the specific Celeron configuration, performance, or maintainer attribution.

Why it matters to Scott

The radar already tracks substantially similar low-memory MoE offloading and expert-streaming claims on “AirLLM — low-VRAM model streaming” and “HotPin — lossless MoE streaming.” The unusually low 2.7GB/Celeron floor bears directly on Scott’s hardware-aware local-inference work, but performance and even the specific benchmark remain independently unverified.
dev:concept.hardware-aware-local-inferenceradar:airllm-low-vram-model-streamingradar:hotpin-lossless-moe-streamingradar:concept.expert-streamingradar:concept.moe-inference
queries asked of Scott's wikis
  • local inference on memory-constrained commodity hardware
  • MoE expert offloading and sparse activation
  • CPU offload versus quantization economics
  • minimum viable hardware for local AI
  • Rust bindings for llama.cpp inference
  • open-model accessibility through low-resource inference

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnMoe expert offloading on a 2-core Celeron with 2.7GB RAMfukoffandie20
🟧 echo.github ⭐The repository demonstrates MoE expert offloading on a two-core Celeron system with 2.7GB of RAM.dhishwasher——

Interpretation history

Decision trace