2026-10-11 17:10 UTC

Independent testing will determine whether the Muse Glimmer deployment on ExecuTorch provides practically fast and reliable on-device agentic inference across supported mobile and edge hardware.

state: expiredheat: lowuncertainty: highknownscott: mediumlocal-inference edge-agents executorchPyTorchExecuTorch

What is this?

Meta introduced Muse Glimmer, an open-weight 30B-parameter model distilled from Muse Spark for on-device agentic workflows, alongside ExecuTorch support for NVIDIA GPUs and Macs with Apple silicon. The supplied sources say it uses roughly 4-bit quantization to fit under 20 GB and a hybrid attention architecture intended to make 128K+ contexts practical on edge hardware; NVIDIA also claims deployment across several GeForce, DGX, and Jetson devices. However, these are primarily Meta, PyTorch, NVIDIA, and model-release claims, and the snippets do not provide identifiable independent benchmarks establishing practical speed, reliability, or broad mobile-hardware support.

Why it matters to Scott

The radar already tracks the same Muse Glimmer local-inference development in `radar:meta-muse-open-weights-local-inference`. The ExecuTorch deployment angle still bears directly on Scott’s hardware-aware local-inference work and Capability Audit/Evaluation-Driven Development position: vendor speed and compatibility claims need representative, trace-backed testing before being treated as production capability.
ip:concept.capability-auditip:concept.evaluation-driven-developmentdev:concept.hardware-aware-local-inferencedev:project.gamepcradar:meta-muse-open-weights-local-inferenceradar:concept.local-inferenceradar:concept.edge-inferenceradar:concept.agent-evaluation
queries asked of Scott's wikis
  • local inference economics and hardware thresholds
  • edge agents versus cloud agent architectures
  • on-device agent reliability and benchmark methodology
  • ExecuTorch deployment projects and constraints
  • long-context KV-cache strategies for local models
  • open-weight models for private agentic workflows

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (6) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnFast, on Device Agentic AI with Muse Glimmer on ExecuTorchandrewstetsenko20
🟧 echo.blog ⭐The primary announcement introducing Muse Glimmer. Meta describes it as “a 30-billion-parameter model optimized for always-on local agent woMeta Superintelligence Labs——
🟠 redditMuse Glimmer was frontier In the model class around 30b models for four days.
LocalLLaMA
InternationalGap3698501172
🟠 redditGitHub - meta-models/meta-oss-cookbook: All recipes for oss models from Meta Inc.
LocalLLaMA
pmttyji322
🟠 redditMuse Glimmer Q8 looping badly/unusable for coding
LocalLLaMA
Electronic_Back1502027
🟠 redditMuse Glimmer 30B with 512k context
LocalLLaMA
mr_il211

Interpretation history

Decision trace