2026-10-11 16:37 UTC

moe-serving

band: warmmomentum: stable score: 0.297
temperature history

Episodes (2)

Fractal-BLTโ€™s publisher claims its released .NET 10 MoE runtime streams weights from NVMe to GPU with zero allocation, potentially enabling local inference on models whose weights exceed GPU memory.
watchingconvergesscott: medium
LocalLLaMA user EmPips reports the Strata runtime streams ~80GB Qwen3.8-Next (IQ3_XXS) from DDR4 to a 7900 XTX at 45-50 tok/s โ€” roughly 2x tuned llama.cpp on the same rig โ€” with similar results reported by users on 12-16GB cards, and independent replication would establish system-RAM weight streaming as a practical route to large-MoE local inference on modest GPUs.
corroboratedconvergesscott: high