2026-10-11 18:01 UTC

moe-inference

band: coolmomentum: stable score: 0.026
temperature history

Episodes (4)

Independent benchmarks will determine whether Meituan’s LongCat-Flash-Lite-Sparse can deliver practical 256K-context inference on 24GB GPUs by combining sparse MoE activation with a RAM-offloaded n-gram lookup table.
expirednovelscott: none
Independent benchmarks will determine whether Slipstream's SSD expert streaming enables 35B–480B MoE coding models to run usefully on memory-constrained Macs without prohibitive latency, swapping, or SSD wear.
expirednovelscott: low
Independent benchmarks will determine whether llama.cpp’s hot-expert GPU cache materially accelerates CPU-offloaded MoE inference on memory-constrained GPUs without regressions across models and quantizations.
expiredconvergesscott: medium
CharacterBumblebee99 claims LayerStoRm's MIT-licensed expert-streaming engine runs 186 GiB GLM-5.3-Flash weights at 24.5 tokens per second at 8K context on 96 GB of GPU VRAM plus roughly 208 GB of pinned host RAM, potentially making oversized MoE models practical on consumer multi-GPU systems.
expiredconvergesscott: medium

Trajectory notes