2026-10-11 17:09 UTC

moe-offloading

band: hotmomentum: stable score: 0.646
temperature history

Episodes (3)

Edge0 claims its released SSD-streaming MoE framework runs its 35B tier at 14.9โ€“17.7 tokens per second on an M4 Pro with 2.9 GiB peak active MLX memory at short contexts, potentially reducing accelerator-memory requirements for local inference without establishing equivalent total-system memory savings.
corroboratedconvergesscott: medium
Atretador claims its released llama.cpp fork fixes expert-cache admission on a 16GB MI50 and raises Qwen3.8-Flash-Next decode throughput from 11.76 to 16.90โ€“17.60 tokens per second at 128K context, potentially accelerating constrained local inference when routing locality supports caching.
watchingknownscott: low
Fork author neuralll claims his released llama.cpp fork's VRAM-filling hot-expert cache (built on csantiago78's PR #27861) roughly doubles decode throughput for GLM-5.3-Flash and MiMo MoE models far larger than total VRAM on two RTX 3090s with unchanged perplexity, and projects further gains per added GPU โ€” if independent multi-GPU users reproduce it, hot-expert caching becomes a practical standard path for memory-constrained local MoE inference.
acceleratingconvergesscott: high