2026-10-11 17:13 UTC

CharacterBumblebee99 claims LayerStoRm's MIT-licensed expert-streaming engine runs 186 GiB GLM-5.3-Flash weights at 24.5 tokens per second at 8K context on 96 GB of GPU VRAM plus roughly 208 GB of pinned host RAM, potentially making oversized MoE models practical on consumer multi-GPU systems.

state: expiredheat: lowuncertainty: highconvergesscott: mediumlocal-inference expert-streaming moe-inference inference-economicsLayerStoRmCharacterBumblebee99

What is this?

The supplied web snippets describe Z.ai’s GLM-5.3-Flash as an MIT-licensed, open-weight mixture-of-experts model with 320B total parameters and 18B active per token, released August 26, 2026. The case attributes to CharacterBumblebee99 a claim that LayerStoRm’s expert-streaming engine runs 186 GiB of quantized weights at 24.5 tokens/second at 8K context using two RTX 5090s, two RTX 5080s and roughly 208 GB of pinned host RAM. None of the search results directly documents LayerStoRm, establishes who develops it, verifies its engine license or independently corroborates that benchmark; the title’s 1M-context claim also lacks supporting measurements here. Several snippets describe much larger GPU-only hosting requirements, but do not evaluate expert streaming, so they do not establish an eight-GPU minimum for the claimed approach.

Why it matters to Scott

LayerStoRm’s claimed expert streaming converges with Scott’s hardware-aware local-inference approach: if reproduced, it could expand the oversized-model serving options worth evaluating for his gamepc substrate, though the hits establish neither sufficient hardware nor runtime compatibility. The benchmark and engine license remain unverified; the radar already tracks related RAM-streaming work in krasis-single-gpu-ornith-397b, but the supplied hits do not show this LayerStoRm development already covered.
dev:concept.hardware-aware-local-inferencedev:project.gamepcradar:concept.expert-streamingradar:concept.moe-inferenceradar:concept.inference-benchmarkingradar:krasis-single-gpu-ornith-397bradar:llama-cpp-hot-expert-gpu-cache
queries asked of Scott's wikis
  • MoE expert streaming CPU offload GPU memory limits
  • local inference economics host RAM multi-GPU hardware
  • open-weight model deployment hardware accessibility
  • coding agent harness local model throughput long context
  • quantized model benchmarking prefill decode performance

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐LayerStoRm open-source expert streaming: 1M context GLM-5.3-Flash [UD-Q4_K_XL] at 24.5 tok/s @8k on just 2× RTX 5090 + 2× RTX 5080 (186 GiB MoE on 96 GB VRAM)
LocalLLaMA
CharacterBumblebee99022

Interpretation history

Decision trace