CharacterBumblebee99 claims LayerStoRm's MIT-licensed expert-streaming engine runs 186 GiB GLM-5.3-Flash weights at 24.5 tokens per second at 8K context on 96 GB of GPU VRAM plus roughly 208 GB of pinned host RAM, potentially making oversized MoE models practical on consumer multi-GPU systems.
state: expiredheat: lowuncertainty: highconvergesscott: mediumlocal-inference expert-streaming moe-inference inference-economicsLayerStoRmCharacterBumblebee99
What is this?
The supplied web snippets describe Z.ai’s GLM-5.3-Flash as an MIT-licensed, open-weight mixture-of-experts model with 320B total parameters and 18B active per token, released August 26, 2026. The case attributes to CharacterBumblebee99 a claim that LayerStoRm’s expert-streaming engine runs 186 GiB of quantized weights at 24.5 tokens/second at 8K context using two RTX 5090s, two RTX 5080s and roughly 208 GB of pinned host RAM. None of the search results directly documents LayerStoRm, establishes who develops it, verifies its engine license or independently corroborates that benchmark; the title’s 1M-context claim also lacks supporting measurements here. Several snippets describe much larger GPU-only hosting requirements, but do not evaluate expert streaming, so they do not establish an eight-GPU minimum for the claimed approach.
Why it matters to Scott
LayerStoRm’s claimed expert streaming converges with Scott’s hardware-aware local-inference approach: if reproduced, it could expand the oversized-model serving options worth evaluating for his gamepc substrate, though the hits establish neither sufficient hardware nor runtime compatibility. The benchmark and engine license remain unverified; the radar already tracks related RAM-streaming work in krasis-single-gpu-ornith-397b, but the supplied hits do not show this LayerStoRm development already covered.
dev:concept.hardware-aware-local-inferencedev:project.gamepcradar:concept.expert-streamingradar:concept.moe-inferenceradar:concept.inference-benchmarkingradar:krasis-single-gpu-ornith-397bradar:llama-cpp-hot-expert-gpu-cache
queries asked of Scott's wikis
- MoE expert streaming CPU offload GPU memory limits
- local inference economics host RAM multi-GPU hardware
- open-weight model deployment hardware accessibility
- coding agent harness local model throughput long context
- quantized model benchmarking prefill decode performance
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-09-07T01:28:41Z
No substantive follow-up has arrived within the observation horizon, leaving LayerStoRm an unverified, hardware-specific expert-streaming lead rather than an established deployment option for Scott. Expiration reflects lack of follow-through, not disproof; accessible code or an independent reproduction would justify reopening.
2026-09-07T01:26:00Z
grounded: converges/medium — LayerStoRm’s claimed expert streaming converges with Scott’s hardware-aware local-inference approach: if reproduced, it could expand the oversized-model serving
2026-09-07T01:23:26Z
case created — A named experimental engine and a concrete hardware-throughput claim establish a distinct local-inference episode, although code availability and measurements remain unverified.
Decision trace
- 09-07 11:28expireNo substantive follow-up has arrived within the observation horizon, leaving LayerStoRm an unverified, hardware-specific expert-streaming lead rather than an established deployment option for Scott. E
- 09-07 11:28alert_silentThe supplied delta contains no new implementation, benchmark evidence, or access change. The original claim still lacks demonstrated source standing or supporting artifacts, so there is no consequenti
- 09-07 11:28alert_routeThe supplied delta contains no new implementation, benchmark evidence, or access change. The original claim still lacks demonstrated source standing or supporting artifacts, so there is no consequenti
- 09-07 11:28alert_silentSubstantive local-inference lead worth retaining for the next briefing, but no time-sensitive decision is apparent. The supplied post reports expert streaming and prompt-cache checkpoints without link
- 09-07 11:28surface_candidateSubstantive local-inference lead worth retaining for the next briefing, but no time-sensitive decision is apparent. The supplied post reports expert streaming and prompt-cache checkpoints without link
- 09-07 11:28alert_routeSubstantive local-inference lead worth retaining for the next briefing, but no time-sensitive decision is apparent. The supplied post reports expert streaming and prompt-cache checkpoints without link
- 09-07 11:26groundLayerStoRm’s claimed expert streaming converges with Scott’s hardware-aware local-inference approach: if reproduced, it could expand the oversized-model serving options worth evaluating for his gamepc
- 09-07 11:23createA named experimental engine and a concrete hardware-throughput claim establish a distinct local-inference episode, although code availability and measurements remain unverified.
- 08-27 22:27expireNo independent benchmark, implementation uptake, or reliability evidence appeared within the observation horizon. This hardware-specific release has faded as a standalone episode, while the broader ex
- 08-27 22:27alert_silentThe only trigger is staleness, with no consequential new evidence or product change; there is nothing Scott needs before the next briefing.
- 08-27 22:27alert_routeThe only trigger is staleness, with no consequential new evidence or product change; there is nothing Scott needs before the next briefing.
- 08-26 02:21sensor_dirtyengagement_update
- 08-25 21:33repriceThe refreshed discussion remains repetitive comparison with adjacent projects and adds no independent benchmark, implementation uptake, or reliability evidence. LayerStoRm is still an unvalidated, har
- 08-25 21:33alert_silentNo consequential new fact emerged; comparative comments without validation or a product change can wait for routine review.
- 08-25 21:33alert_routeNo consequential new fact emerged; comparative comments without validation or a product change can wait for routine review.
- 08-25 21:21sensor_dirtycomment_update
- 08-25 18:30repriceThe refreshed discussion only identifies overlap with FreeToken, Colibri, and automated tensor scheduling; it adds no benchmark, adoption, or implementation evidence. LayerStoRm remains a hardware-spe
- 08-25 18:30alert_silentThe new comments are comparative discussion rather than validation or a consequential product change, so there is nothing Scott needs before the next briefing.
- 08-25 18:30alert_routeThe new comments are comparative discussion rather than validation or a consequential product change, so there is nothing Scott needs before the next briefing.
- 08-25 18:21sensor_dirtycomment_update
- 08-25 14:27repriceThe reobservation adds no benchmark, implementation uptake, or developer evidence, so LayerStoRm remains an unvalidated instance of an already tracked expert-streaming pattern. Cool the case pending i
- 08-25 14:27alert_silentNothing consequential changed since the early-release post; unchanged engagement and the absence of independent validation can wait for routine review.
- 08-25 14:27alert_routeNothing consequential changed since the early-release post; unchanged engagement and the absence of independent validation can wait for routine review.
- 08-25 14:26alert_silentA low-engagement Reddit post describes an experimental, highly hardware-specific implementation and self-reported throughput, but provides no visible repository or independent benchmark. It is relevan
- 08-25 14:26surface_candidateA low-engagement Reddit post describes an experimental, highly hardware-specific implementation and self-reported throughput, but provides no visible repository or independent benchmark. It is relevan
- 08-25 14:26alert_routeA low-engagement Reddit post describes an experimental, highly hardware-specific implementation and self-reported throughput, but provides no visible repository or independent benchmark. It is relevan
- 08-25 14:26groundThe radar already tracks this same expert-streaming/offloading question in “AirLLM low-VRAM model streaming,” “ExpertCache GPT-OSS 120B M1,” and “llama.cpp hot-expert GPU cache.” LayerStoRm is another
- 08-25 14:24createThe first-party early release describes a concrete expert-streaming implementation for severely VRAM-constrained multi-GPU inference.