The MoE Offload Bench maintainer claims the released implementation can offload sparse-model experts on a two-core Celeron with 2.7GB of RAM, potentially extending local MoE inference to extremely constrained commodity systems.
state: expiredheat: lowuncertainty: highknownscott: mediumlocal-inference inference-economics open-modelsdhishwasher
What is this?
The case reports a released MoE expert-offloading implementation whose maintainer, identified as dhishwasher, claims it can run on a two-core Celeron with 2.7GB of RAM. The returned snippets substantiate the underlying approach: llama.cpp can place MoE expert tensors in CPU RAM, while sparse models activate only a subset of experts per token; related releases add Rust bindings and packaged model integrations. However, the snippets do not independently identify the claimed benchmark repository or verify the specific Celeron configuration, performance, or maintainer attribution.
Why it matters to Scott
The radar already tracks substantially similar low-memory MoE offloading and expert-streaming claims on “AirLLM — low-VRAM model streaming” and “HotPin — lossless MoE streaming.” The unusually low 2.7GB/Celeron floor bears directly on Scott’s hardware-aware local-inference work, but performance and even the specific benchmark remain independently unverified.
dev:concept.hardware-aware-local-inferenceradar:airllm-low-vram-model-streamingradar:hotpin-lossless-moe-streamingradar:concept.expert-streamingradar:concept.moe-inference
queries asked of Scott's wikis
- local inference on memory-constrained commodity hardware
- MoE expert offloading and sparse activation
- CPU offload versus quantization economics
- minimum viable hardware for local AI
- Rust bindings for llama.cpp inference
- open-model accessibility through low-resource inference
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-09-04T01:25:52Z
No verification, benchmark detail, or independent reproduction emerged within the initial window, leaving the unusually low hardware-floor claim as an unsubstantiated overlap with existing expert-streaming work.
2026-09-02T00:35:52Z
No new evidence verifies the repository’s specific hardware floor, model, throughput, or reproducibility, so the claim remains an interesting but overlapping implementation lead rather than a demonstrated advance.
2026-09-02T00:32:07Z
grounded: known/medium — The radar already tracks substantially similar low-memory MoE offloading and expert-streaming claims on “AirLLM — low-VRAM model streaming” and “HotPin — lossle
2026-09-02T00:30:17Z
case created — The repository is a concrete and reproducible local-inference artifact exploring a notable hardware constraint.
Decision trace
- 09-04 11:25expireNo verification, benchmark detail, or independent reproduction emerged within the initial window, leaving the unusually low hardware-floor claim as an unsubstantiated overlap with existing expert-stre
- 09-04 11:25alert_silentThe staleness trigger carries no new consequential evidence; the case can fade unless a reproducible benchmark or technical implementation detail appears later.
- 09-04 11:25alert_routeThe staleness trigger carries no new consequential evidence; the case can fade unless a reproducible benchmark or technical implementation detail appears later.
- 09-02 10:35repriceNo new evidence verifies the repository’s specific hardware floor, model, throughput, or reproducibility, so the claim remains an interesting but overlapping implementation lead rather than a demonstr
- 09-02 10:35alert_silentThe reobservation is unchanged and adds no consequential delta; technical inspection or independent reproduction can wait for the normal briefing.
- 09-02 10:35alert_routeThe reobservation is unchanged and adds no consequential delta; technical inspection or independent reproduction can wait for the normal briefing.
- 09-02 10:32alert_silentA public repository now claims MoE expert offloading on a two-core Celeron with 2.7GB RAM, but the supplied evidence provides no benchmark, model, throughput, implementation details, or independent re
- 09-02 10:32surface_candidateA public repository now claims MoE expert offloading on a two-core Celeron with 2.7GB RAM, but the supplied evidence provides no benchmark, model, throughput, implementation details, or independent re
- 09-02 10:32alert_routeA public repository now claims MoE expert offloading on a two-core Celeron with 2.7GB RAM, but the supplied evidence provides no benchmark, model, throughput, implementation details, or independent re
- 09-02 10:32groundThe radar already tracks substantially similar low-memory MoE offloading and expert-streaming claims on “AirLLM — low-VRAM model streaming” and “HotPin — lossless MoE streaming.” The unusually low 2.7
- 09-02 10:30createThe repository is a concrete and reproducible local-inference artifact exploring a notable hardware constraint.