2026-10-11 17:11 UTC

Independent testing will determine whether Maple-Preview’s ternary 20B MoE sustains roughly 120 tokens per second on an iPhone while retaining practically useful model quality.

state: expiredheat: lowuncertainty: highknownscott: lowlocal-inference ternary-models edge-aiDeepGrove

What is this?

Maple-Preview is presented in the case as a ternary 20B mixture-of-experts model from DeepGrove, with a claimed generation speed of roughly 120 tokens per second on an iPhone. One supplied video snippet reports about 119 tokens per second for an unnamed ternary model using a llama.cpp interface, but it does not clearly identify the device, Maple-Preview, DeepGrove, or assess model quality. The other results concern unrelated models or general local-inference performance, so the supplied snippets do not yet establish independent confirmation of Maple-Preview’s speed or practical quality.

Why it matters to Scott

The case adds no new position beyond Scott’s Capability Audit and hardware-aware local inference work: claimed mobile throughput still needs representative, independent quality and performance testing. The radar already tracks nearly identical validation questions in “Swiftlet ultralow-memory inference” and “Bonsai extreme quantization”; absent independent results, Maple-Preview is another unverified instance rather than a development that changes Scott’s view or practice.
ip:concept.capability-auditdev:concept.hardware-aware-local-inferenceradar:swiftlet-ultralow-memory-inferenceradar:bonsai-extreme-quantization
queries asked of Scott's wikis
  • ternary model quality versus quantization
  • Apple Silicon local inference economics
  • on-device MoE routing and performance
  • independent evaluation of local LLM claims
  • edge AI privacy and model sovereignty
  • llama.cpp mobile inference optimization

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhoneedwardbzhang17152
🟧 echo.blog ⭐Maple-Preview is presented as a ternary 20B mixture-of-experts model running at roughly 120 tokens per second on an iPhone.DeepGrove——
🟠 redditMaple-Preview: 20B-A1B ternary-weight reasoning open-weight LLM
LocalLLaMA
cafedude13655
🟠 redditI've added Maple-Preview to Mference, got 40 tps generation with 500MB of used RAM on Air M4
LocalLLaMA
netikas35
🟠 redditIs ternary (1.58-bit) LLMs making a come back?
LocalLLaMA
Individual-Dot54882117

Interpretation history

Decision trace