2026-10-11 16:38 UTC

Independent reproduction will determine whether ToMoE can convert dense LLMs into sparse mixture-of-experts models that materially reduce active inference compute without unacceptable quality loss.

state: watchingheat: lowuncertainty: highconvergesscott: mediummixture-of-experts model-sparsity local-inference

What is this?

ToMoE is a research method by Shangqian Gao et al. (TMLR 2026, ICML 2026 poster, arXiv 2501.15316, official code at github.com/gaosh/ToMoE) that converts dense LLMs into sparse mixture-of-experts models by differentiable dynamic structural pruning — splitting MLP layers into experts with a fixed per-token active-parameter budget, 'without requiring any weight updates,' i.e. no retraining. It sits in a growing 2025–26 family of training-free or cheap dense→MoE conversion methods (CMoE, Dense2MoE, ExpertWeaver, DOT-MoE). The case's own evidence adds that a v2 paper now circulates and that a community thread reading its benchmark tables reports severe quality loss (MMLU 67.22→36.31, explicitly hedged), contradicting the near-lossless framing — the web snippets supplied here neither mention v2 nor that critique, and confirm only the v1 claims and the authors' own implementation, not any independent reproduction or released converted checkpoints.

Why it matters to Scott

The v2 quality challenge was surfaced not by a reproduction lab but by a commenter doing the one thing Scott's canon demands — reading the paper's own tables as verbatim exhibits, hedged with 'if I am reading it correctly' — a dated receipt for witness-not-oracle and the evidence-class ladder: a near-lossless claim defended only at announcement class gets caught by its own benchmark tables, with the exit falsifier supplied post hoc exactly as the falsifiability spine predicts when none was pre-registered. Substantively it closes off, negatively, a potential dense→MoE conversion lever for cheaper local serving on his gamepc/Ollama stack — the technique now shows severe quality loss at usable budgets AND lacks runtime support — so the actionable position is to keep the case open only for converted-checkpoint releases and to apply the same lowered prior to siblings (Prox, Dense2MoE, CMoE).
ip:source.witness-not-oracle-ebookip:concept.evidence-class-ladderip:framework.falsifiability-spinedev:concept.hardware-aware-local-inferenceradar:concept.reproducibilityradar:concept.benchmark-integrityradar:quantization-aware-healing-validationradar:haar-wavelet-llm-pruningradar:late-layer-moe-expert-expansionradar:glm-lossless-weight-compression
queries asked of Scott's wikis
  • local inference runtime support for converted MoE checkpoints (llama.cpp / vLLM / expert sharding)
  • MoE memory footprint vs FLOP savings on consumer GPUs at batch-1 — does active-parameter reduction actually speed local serving
  • structural pruning quality-loss tradeoff positions and thresholds
  • training-free dense-to-sparse conversion methods — evaluation and reproducibility standards
  • quantization vs sparsity economics for local model deployment
  • paper-claims-vs-reproduction policy on compression/efficiency research

Measured heat

now 7 pts/hpeak 39 pts/hcomments 2/hpeers p81momentum: steady1 platformsage 1154h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

08-24 13:54⭐ origin directly observed[Paper] ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning
pmttyji on r/LocalLLaMA
—
09-28 18:35first on r/LocalLLaMA · published · +844.7hWhat do you think about ToMoE v2 paper, converting dense model to MoE model at near lossless accuracy? I feel Qwen3.8-27B-A16B or something along those lines would be amazing, though there are architectural hurdles, as well as need folr training data.
jinnyjuice
—
08-24 13:54amplified on r/LocalLLaMA 👑reddit.post.1vx3img
pmttyji
peak 270 · 29 comments · 60% of case engagement
09-28 18:35amplified on r/LocalLLaMAreddit.post.1wsmsnx
jinnyjuice
peak 73 · 12 comments · 17% of case engagement
10-11 05:45amplified on r/LocalLLaMAreddit.post.1x3078i
Bayliner1980
peak 84 · 33 comments · 23% of case engagement
08-24 14:20our radar first saw it · +0.4hdiscovery anchor: reddit.post.1vx3img—

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐[Paper] ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning
LocalLLaMA
pmttyji26828
🟠 redditWhat do you think about ToMoE v2 paper, converting dense model to MoE model at near lossless accuracy? I feel Qwen3.8-27B-A16B or something along those lines would be amazing, though there are architectural hurdles, as well as need folr training data.
LocalLLaMA
jinnyjuice7312
🟠 redditConverting dense models into Mixture-of-Experts
LocalLLaMA
Bayliner19808433

Interpretation history

Decision trace