Independent benchmarks will determine whether Meta’s released MobileMoE models establish a superior quality-efficiency tradeoff among sub-3GB on-device language models.
state: expiredheat: lowuncertainty: highconvergesscott: mediummobile-inference open-models mixture-of-expertsMeta
What is this?
MobileMoE is a family of on-device mixture-of-experts language models presented by Meta researchers, with 0.3–0.9B active parameters and 1.3–5.3B total parameters, designed around mobile memory and compute constraints. The paper reports that the models outperform larger dense or MoE baselines while using fewer active parameters; one report cites 1.49GB peak memory on a Galaxy S25 and up to 3.8× faster execution on an iPhone 16 Pro. These claims currently rest mainly on the authors’ benchmarks and secondary reporting, so the asserted quality-efficiency frontier still needs independent validation.
Why it matters to Scott
Meta’s claimed sub-3GB quality-efficiency frontier converges with Scott’s Model Barbell and Usable Mass positions: cheap, deployable models can be more valuable than larger but impractical capability, especially when hardware constraints are explicit. Independent results could affect his local-inference and model-routing choices, but until those results exist this is a promising candidate rather than a demonstrated shift.
ip:concept.model-barbellip:concept.usable-mass-over-unusable-powerdev:concept.hardware-aware-local-inferenceip:concept.evaluation-driven-developmentradar:concept.on-device-airadar:concept.small-modelsradar:concept.mixture-of-expertsradar:concept.benchmark-integrity
queries asked of Scott's wikis
- on-device inference economics and privacy
- small-model quality versus capability tradeoffs
- mixture-of-experts under memory constraints
- open-weight models for local AI systems
- benchmark trust and independent model evaluation
- local model routing and agent architectures
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-08-28T20:42:25Z
The validation window has faded without an independent benchmark, device implementation, or artifact-access change, so this episode has not progressed beyond an unvalidated release. A substantive test result can reopen it as a new episode.
2026-08-26T20:36:25Z
No independent benchmark, device implementation, or artifact-access update has appeared after 48 hours, so the case remains an unvalidated release rather than evidence of a new quality-efficiency frontier. The initial attention has cooled, though later testing could still revive it.
2026-08-24T19:58:24Z
The refreshed discussion remains anticipatory: users intend to test the models and note missing SFT/QAT artifacts, but no independent benchmark or implementation result yet validates the claimed quality-efficiency frontier.
2026-08-24T19:34:06Z
grounded: converges/medium — Meta’s claimed sub-3GB quality-efficiency frontier converges with Scott’s Model Barbell and Usable Mass positions: cheap, deployable models can be more valuable
2026-08-24T19:30:26Z
case created — Meta has released concrete open model artifacts for memory-constrained devices, while their comparative quality and efficiency remain unvalidated.
Decision trace
- 08-29 06:42expireThe validation window has faded without an independent benchmark, device implementation, or artifact-access change, so this episode has not progressed beyond an unvalidated release. A substantive test
- 08-29 06:42alert_silentThere is no new consequential fact—only continued absence of validation—and elapsed silence is not alert-worthy.
- 08-29 06:42alert_routeThere is no new consequential fact—only continued absence of validation—and elapsed silence is not alert-worthy.
- 08-27 06:36repriceNo independent benchmark, device implementation, or artifact-access update has appeared after 48 hours, so the case remains an unvalidated release rather than evidence of a new quality-efficiency fron
- 08-27 06:36alert_silentThe only delta is elapsed time without new validation; the release was already routed, and there is no consequential new fact for Scott today.
- 08-27 06:36alert_routeThe only delta is elapsed time without new validation; the release was already routed, and there is no consequential new fact for Scott today.
- 08-26 00:21sensor_dirtyengagement_update
- 08-25 07:21sensor_dirtyengagement_update
- 08-25 05:58repriceThe refreshed discussion remains anticipatory: users intend to test the models and note missing SFT/QAT artifacts, but no independent benchmark or implementation result yet validates the claimed quali
- 08-25 05:58alert_silentThe established release was already routed; the new delta is only modest engagement and repeated plans to test, with no independent results or material access change to report.
- 08-25 05:58alert_routeThe established release was already routed; the new delta is only modest engagement and repeated plans to test, with no independent results or material access change to report.
- 08-25 05:52alert_shadowThe first-party Hugging Face collection establishes a practical new on-device MoE release with 0.3B–0.9B active parameters and advertised INT4 footprints below 3 GB, directly relevant to local inferen
- 08-25 05:52alert_routeThe first-party Hugging Face collection establishes a practical new on-device MoE release with 0.3B–0.9B active parameters and advertised INT4 footprints below 3 GB, directly relevant to local inferen
- 08-25 05:34groundMeta’s claimed sub-3GB quality-efficiency frontier converges with Scott’s Model Barbell and Usable Mass positions: cheap, deployable models can be more valuable than larger but impractical capability,
- 08-25 05:30createMeta has released concrete open model artifacts for memory-constrained devices, while their comparative quality and efficiency remain unvalidated.