2026-10-11 17:16 UTC

llama.cpp contributor predatar claims PR #28086 raises IQ3-quantized MoE decode throughput on Apple Silicon Metal from about 65.6 to 73.9 tokens per second, potentially improving local sparse-model inference if merged.

state: watchingheat: lowuncertainty: highknownscott: lowlocal-inference llama-cpp apple-silicon inference-optimizationpredatarllama.cpp

What is this?

The supplied search answer reports that a proposed llama.cpp Metal optimization improves IQ3-quantized mixture-of-experts decode throughput on Apple Silicon from 65.6 to 73.9 tokens per second—about a 13% gain. The surrounding snippets establish that llama.cpp and MLX are competing local-inference paths on Apple Silicon and that performance varies substantially by model, quantization, workload, and hardware. However, the snippets do not independently verify predatar’s attribution or PR #28086, and the evidence title instead names PR #13388, “metal: optimize MoE for large batches,” so the exact primary artifact and merge status remain unclear.

Why it matters to Scott

This is another implementation-level example of Scott’s “Hardware-aware local inference” position, and the radar already tracks llama.cpp, Apple Silicon, MoE inference, and inference optimization. The claimed roughly 13% gain could become operationally useful, but the conflicting PR identifiers, unclear merge status, and lack of independent validation mean it does not yet change what Scott should build or argue.
dev:concept.hardware-aware-local-inferenceradar:concept.llama-cppradar:concept.apple-silicon-inferenceradar:concept.moe-inferenceradar:concept.inference-optimization
queries asked of Scott's wikis
  • local inference economics on Apple Silicon
  • GGUF and IQ quantization trade-offs
  • MoE inference optimization for local models
  • llama.cpp versus MLX backend strategy
  • benchmark methodology for local inference
  • Apple Silicon local agent deployment

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 12530h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

05-07 14:00⭐ origin echo-reconstructedThe original primary artifact is llama.cpp PR #13388, “metal : optimize MoE for large batches.” It describes remapping per-expert inputs, pe
Georgi Gerganov (ggerganov) on github (echo) · attributed from reddit.post.1w46a73
—
09-01 08:55first on r/LocalLLaMA · published · +11562.9hLlama cpp metal moe optimization
predatar
—
09-01 08:55amplified on r/LocalLLaMA 👑reddit.post.1w46a73
predatar
peak 6 · 7 comments · 69% of case engagement
09-03 14:13amplified on r/LocalLLaMAreddit.post.1w68lyj
predatar
peak 4 · 2 comments · 32% of case engagement
09-01 09:20our radar first saw it · +11563.3hdiscovery anchor: reddit.post.1w46a73—

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditLlama cpp metal moe optimization
LocalLLaMA
predatar67
🟧 echo.github ⭐The original primary artifact is llama.cpp PR #13388, “metal : optimize MoE for large batches.” It describes remapping per-expert inputs, peGeorgi Gerganov (ggerganov)——
🟠 redditllama cpp Metal moe improvements for decode + prefill + caching (tested on Qwen3-30B-A3B), looking for Qwen3.8-Flash-Next testers
LocalLLaMA
predatar42

Interpretation history

Decision trace