2026-10-11 17:15 UTC

mixture-of-experts

band: warmmomentum: stable score: 0.523
temperature history

Episodes (12)

Independent reproduction will determine whether the published GGUF-based workflow can LoRA-train Qwen3.6-35B-A3B within 16GB of VRAM using APEX quantization and fused kernels without prohibitive performance or quality tradeoffs.
expiredknownscott: medium
Independent evaluations will determine whether Kwaipilot’s 35B-total, 3B-active KAT-Coder-V2.5-Dev delivers competitive agentic-coding and tool-use performance among similarly sized open-weight models.
resolvedknownscott: medium
Independent testing will determine whether Lumabri can practically distribute storage and inference for very large MoE models across ordinary networked computers.
expiredconvergesscott: medium
Independent reproduction will determine whether optimal-transport methods materially reduce mixture-of-experts load imbalance and improve training efficiency over existing balancing techniques.
expiredknownscott: low
The live open training run of a 535B-parameter, 23B-active mixture-of-experts model will publish usable checkpoints, training details, and results sufficient for outside scrutiny of the training process.
watchingconvergesscott: medium
Independent reproduction will determine whether ToMoE can convert dense LLMs into sparse mixture-of-experts models that materially reduce active inference compute without unacceptable quality loss.
watchingconvergesscott: medium
Independent benchmarks will determine whether Meta’s released MobileMoE models establish a superior quality-efficiency tradeoff among sub-3GB on-device language models.
expiredconvergesscott: medium
CharacterBumblebee99 claims LayerStoRm's MIT-licensed expert-streaming engine runs 186 GiB GLM-5.3-Flash weights at 24.5 tokens per second at 8K context on 96 GB of GPU VRAM plus roughly 208 GB of pinned host RAM, potentially making oversized MoE models practical on consumer multi-GPU systems.
expiredconvergesscott: medium
IFM claims its released K2 Horizon family combines competitive model quality, multiple open local-inference sizes, and a sparse 36B model with 4B active parameters, potentially lowering the cost and improving the reproducibility of capable local deployments.
corroboratedconvergesscott: medium
The paper’s authors claim increasing expert activation only in the later layers of Qwen sparse-MoE models reduces reasoning-token use by about 8.5% without retraining or material quality loss, potentially lowering inference cost through a runtime-only change.
watchingconvergesscott: medium
IQuest claims its released IQuest-Q1 β€” a 320B-total/~15B-active MoE purpose-built for agentic coding, reasoning, and multi-step tool use β€” is a capable open-weight coding-agent model; community adoption and independent measurement of it in local coding-agent workflows will determine whether it earns a practical place or fades as another unreleased-in-practice announcement.
watchingconvergesscott: medium
Prism ML (SkyIsNotGreen) claims its released Scion-35B-A3B β€” a 35B-A3B MoE shipped as one 11.3GB GGUF with ternary PQ2_0 expert banks plus embedded trained corrections at 2.61 bpw, a bundled MTP drafter, and a required llama.cpp fork β€” achieves Q4-class task retention at roughly half Q4_K_M's size, making ternary MoE a practically servable local tier; independent adoption and reproduced benchmarks confirm it, quiet fade closes it.
seednovelscott: medium

Trajectory notes