2026-10-11 16:38 UTC

IngeniousIdiocy claims their published ds4 branch runs GLM-5.3 Flash Q4 on an M3 Ultra at over 38 output tokens per second in a roughly 200K-context Claude Code workload, potentially making long-context local coding more responsive on Apple hardware.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: mediumlocal-inference inference-optimization coding-agentsIngeniousIdiocy

What is this?

GLM-5.3-Flash is an MIT-licensed open-weight mixture-of-experts model (320B total / 18B active, 1M-token context via hybrid sparse+linear attention, released Aug 26 2026) positioned as a budget agentic-coding model. Reddit user IngeniousIdiocy reports that a custom 'ds4' fork they published sustains >38 output tok/s running GLM-5.3-Flash Q4 on an M3 Ultra inside ~200K-context Claude Code sessions, with the thread describing single-stream decode at ~59% of the machine's measured memory bandwidth. The supplied results do not independently verify those Apple-specific numbers โ€” they trace back to the author's own posts โ€” and other Mac users report default stacks slowing sharply past 50โ€“60k context (Unsloth's llama.cpp/Metal docs show the same long-context decode decay), so the claim rests on the ds4 implementation itself. Independent runs on other hardware (2x DGX Sparks via vLLM at ~40 tok/s decode at 100k per the case; a ~33 tok/s local YouTube test) and a third-party M3 Ultra DS4/Claude-Code benchmark (echalupa.com) make the broader long-context-local-coding pattern credible while leaving the specific M3 Ultra figures author-reported.

Why it matters to Scott

The -dysangel- replication (different author, hardware and stack: 2x DGX Sparks, published vLLM TP2 recipe, ~40 tok/s decode at 100k in 260k-context Claude Code) upgrades this from lone-author claim to corroborated cross-hardware pattern, and the two inspectable recipes now bear directly on the local-backend option in Ask โ€” though the Apple-specific numbers stay author-reported, coding quality is unvalidated, and neither demonstrated platform matches Scott's gamepc/Mac hardware, so this informs rather than changes that choice. The implementations also independently arrive at his hardware-aware local-inference methodology (Q4 MoE, MTP, explicit placement against the memory-bandwidth ceiling), giving him dated receipts in a territory his dev canon already occupies.
dev:project.askdev:concept.hardware-aware-local-inferencedev:project.gamepcradar:concept.local-inferenceradar:concept.apple-silicon-inferenceradar:concept.memory-bandwidthradar:concept.speculative-decodingradar:concept.mixture-of-expertsradar:concept.quantizationradar:concept.long-contextradar:concept.coding-agentsradar:mlxfast-agent-engine-rewriteradar:apple-m5-ultra-local-inference
queries asked of Scott's wikis
  • Ask project local model backend option โ€” serving a local LLM behind a coding agent
  • Apple Silicon unified-memory bandwidth as the ceiling for local LLM decode speed
  • long-context decode decay โ€” KV cache, prefix caching and cache-hit rates in agent loops
  • speculative decoding / multi-token prediction (MTP) for local inference speedups
  • MoE few-active-parameter models at Q4 โ€” local deployment tradeoffs and quantization choices
  • local vs API coding-agent backend โ€” cost, privacy and hardware break-even economics

Measured heat

now 0 pts/hpeak 80 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 771h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-09 13:24 (minted)โญ origin echo-reconstructedThe author links this ds4 branch as the implementation behind reported GLM-5.3 Flash Q4 throughput on M3 Ultra, distinguishing over 38 token
IngeniousIdiocy on github (echo) ยท attributed from reddit.post.1wbkpnw ยท published time unknown
โ€”
09-09 12:51first on r/LocalLLaMA ยท published ยท lag ?GLM 5.3 Flash Q4 @ 60tps / 550tps on M3 Ultra
IngeniousIdiocy
โ€”
09-09 12:51amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wbkpnw
IngeniousIdiocy
peak 131 ยท 38 comments ยท 30% of case engagement
09-15 11:51amplified on r/LocalLLaMAreddit.post.1wgy6tm
IngeniousIdiocy
peak 44 ยท 21 comments ยท 11% of case engagement
10-04 20:00amplified on r/LocalLLaMAreddit.post.1wxp21n
-dysangel-
peak 58 ยท 22 comments ยท 14% of case engagement
10-04 21:53amplified on r/LocalLLaMAreddit.post.1wxrozq
swiebertjee
peak 54 ยท 29 comments ยท 15% of case engagement
10-05 13:27amplified on r/LocalLLaMAreddit.post.1wy8k3m
cryotic
peak 121 ยท 35 comments ยท 27% of case engagement
10-08 15:56amplified on r/LocalLLaMAreddit.post.1x0unn5
challis88ocarina
peak 6 ยท 9 comments ยท 3% of case engagement
09-09 13:20our radar first saw it ยท lag ?discovery anchor: reddit.post.1wbkpnwโ€”
pace: p84 vs 519 stories at the 720h mark (now 771h old) โ€” ahead of mercury-25-diffusion-inference (1.0x), behind memctl-versioned-agent-memory (1.0x)

Evidence (7) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditGLM 5.3 Flash Q4 @ 60tps / 550tps on M3 Ultra
LocalLLaMA
IngeniousIdiocy13138
๐ŸŸง echo.github โญThe author links this ds4 branch as the implementation behind reported GLM-5.3 Flash Q4 throughput on M3 Ultra, distinguishing over 38 tokenIngeniousIdiocyโ€”โ€”
๐ŸŸ  redditDeepSeek V4.1F Q4 on M3 Ultra with native DSpark MTP (40tps / 800tps)
LocalLLaMA
IngeniousIdiocy4221
๐ŸŸ  redditFully local little parkour sim
LocalLLaMA
-dysangel-5822
๐ŸŸ  redditFor dual DGX spark users; GLM 5.3 flash got a 50%+ performance boost
LocalLLaMA
swiebertjee5129
๐ŸŸ  redditM5 Ultra 256 running GLM 5.3 Flash 68.8 tok/s
LocalLLaMA
cryotic12135
๐ŸŸ  redditMTP in llama.cpp now decodes competitively with ds4 using GLM 5.3 Flash
LocalLLaMA
challis88ocarina59

Interpretation history

Decision trace