GLM-5.3-Flash is an MIT-licensed open-weight mixture-of-experts model (320B total / 18B active, 1M-token context via hybrid sparse+linear attention, released Aug 26 2026) positioned as a budget agentic-coding model. Reddit user IngeniousIdiocy reports that a custom 'ds4' fork they published sustains >38 output tok/s running GLM-5.3-Flash Q4 on an M3 Ultra inside ~200K-context Claude Code sessions, with the thread describing single-stream decode at ~59% of the machine's measured memory bandwidth. The supplied results do not independently verify those Apple-specific numbers โ they trace back to the author's own posts โ and other Mac users report default stacks slowing sharply past 50โ60k context (Unsloth's llama.cpp/Metal docs show the same long-context decode decay), so the claim rests on the ds4 implementation itself. Independent runs on other hardware (2x DGX Sparks via vLLM at ~40 tok/s decode at 100k per the case; a ~33 tok/s local YouTube test) and a third-party M3 Ultra DS4/Claude-Code benchmark (echalupa.com) make the broader long-context-local-coding pattern credible while leaving the specific M3 Ultra figures author-reported.
The -dysangel- replication (different author, hardware and stack: 2x DGX Sparks, published vLLM TP2 recipe, ~40 tok/s decode at 100k in 260k-context Claude Code) upgrades this from lone-author claim to corroborated cross-hardware pattern, and the two inspectable recipes now bear directly on the local-backend option in Ask โ though the Apple-specific numbers stay author-reported, coding quality is unvalidated, and neither demonstrated platform matches Scott's gamepc/Mac hardware, so this informs rather than changes that choice. The implementations also independently arrive at his hardware-aware local-inference methodology (Q4 MoE, MTP, explicit placement against the memory-bandwidth ceiling), giving him dated receipts in a territory his dev canon already occupies.
dev:project.askdev:concept.hardware-aware-local-inferencedev:project.gamepcradar:concept.local-inferenceradar:concept.apple-silicon-inferenceradar:concept.memory-bandwidthradar:concept.speculative-decodingradar:concept.mixture-of-expertsradar:concept.quantizationradar:concept.long-contextradar:concept.coding-agentsradar:mlxfast-agent-engine-rewriteradar:apple-m5-ultra-local-inference
queries asked of Scott's wikis
- Ask project local model backend option โ serving a local LLM behind a coding agent
- Apple Silicon unified-memory bandwidth as the ceiling for local LLM decode speed
- long-context decode decay โ KV cache, prefix caching and cache-hit rates in agent loops
- speculative decoding / multi-token prediction (MTP) for local inference speedups
- MoE few-active-parameter models at Q4 โ local deployment tradeoffs and quantization choices
- local vs API coding-agent backend โ cost, privacy and hardware break-even economics
2026-10-09T11:54:41Z
New independent evidence: challis88ocarina reports MTP in llama.cpp now decodes competitively with ds4 on GLM-5.3 Flash (Apple Silicon), adding a fourth stack (llama.cpp) to the corroborated cross-hardware pattern. The M3 Ultra long-context figures remain author-reported only; the new datapoint doesn't specify context length or workload. Episode velocity is fully cooled (0 pts/h, 30 days old) but the topic neighbourhood stays hot.
2026-10-08T23:06:43Z
evidence attached: reddit.post.1x0unn5 โ User reports MTP in llama.cpp now decoding competitively with ds4 branch on GLM 5.3 Flash, providing independent performance corroboration for the case's Apple Silicon optimization claims.
2026-10-08T02:40:13Z
The cryotic post's velocity spike has fully decayed (0 pts/h now vs 12.5 peak) and the episode's measured heat is cooling at 28 days old; the cross-hardware pattern stands corroborated across four stacks but the Apple-specific long-context M3 Ultra figures remain author-reported only, and the independent M5 Ultra datapoint is short-context without workload spec. Heat drops to low on settled evidence, not fading interest โ the topic neighbourhood stays hot.
2026-10-05T16:07:23Z
The cryotic post adds a genuinely independent Apple-silicon line (M5 Ultra 256, oMLX 0.7.0 oQ4e+MTP, 68.8 tok/s) โ a fourth stack and third hardware vendor โ but its comment section contextualizes rather than replicates: it's a short-context benchmark with no workload spec (swiebertjee calls it meaningless without one), and per oMLX's own benchmarks base 4-bit is already 60-70 tok/s short-context, so it supports general Apple-Ultra viability without verifying the ds4/M3-Ultra long-context figures, which stay author-reported. Heat lifts to medium on a fresh top-decile peer mover landing directly in Scott's Apple-inference territory, not on episode velocity, which is still cooling.
2026-10-05T15:30:48Z
evidence attached: reddit.post.1wy8k3m โ Independent, high-engagement data point running GLM-5.3 Flash faster (68.8 tok/s, oMLX+MTP) on Apple Ultra hardware than the ds4/M3 Ultra claim, materially contextualizing the case's viability thesis.
2026-10-05T14:31:39Z
Follow-ups are consolidation, not expansion: -dysangel- vouches swiebertjee's boosted recipe is the same lineage as the parkour-sim vLLM build (the DGX Spark cluster is one tight-knit line rather than a third fully independent builder), a commenter points to a fourth EXL3 repo, and the transient 13.5x velocity spike decayed to <1 pt/h of question-and-banter discussion. The cross-hardware pattern stands established while the Apple-specific numbers stay author-reported, so the case settles into corroborated-and-cool with heat dropping to low.
2026-10-04T22:48:26Z
A third independent builder (swiebertjee) published a 50โ90% decode-boost recipe for GLM 5.3 Flash on dual DGX Spark and claims the tuned GLM now beats DeepSeek v4.1 Flash โ the pattern has crossed from one-off demonstrations to active community tuning, further entrenching establishment. The Apple-specific M3 Ultra numbers remain author-reported, and per-post engagement is shrinking; medium heat rides on the expanding periphery (three builders, three published recipes), not on velocity.
2026-10-04T22:25:54Z
evidence attached: reddit.post.1wxrozq โ Extends the GLM-5.3-Flash local-performance episode to dual DGX Spark with a published 50-90% decode recipe, claiming it beats DeepSeek v4.1 Flash.
2026-10-04T21:41:04Z
grounded: converges/medium โ The -dysangel- replication (different author, hardware and stack: 2x DGX Sparks, published vLLM TP2 recipe, ~40 tok/s decode at 100k in 260k-context Claude Code
2026-10-04T21:28:43Z
An unrelated builder (-dysangel-) reproduced the long-context local GLM-5.3 Flash coding pattern on different hardware and a different stack (2x DGX Sparks, published vLLM TP2 recipe, ~40 tok/s decode at 100k in 260k-context Claude Code), giving the case its second independent evidence line and moving it from lone-author claim to corroborated pattern. The specific M3 Ultra/ds4 numbers stay author-reported and coding quality remains unvalidated, but the case now carries inspectable recipes that bear on Scott's local-backend option for Ask, lifting its relevance.
2026-10-04T20:26:26Z
evidence attached: reddit.post.1wxp21n โ Independent cross-hardware replication of the GLM-5.3 Flash long-context local-coding pattern: ~40 tok/s decode at 100k and 260k-context Claude Code on 2x DGX Sparks with a published vLLM TP2 recipe.
2026-09-15T12:23:13Z
The same author now reports a sustained DeepSeek V4.1 Flash agent run on a 512GB M3 Ultra, adding a concrete follow-on implementation result rather than another throughput comparison. This supports continued ds4 optimization work, but neither independently corroborates the original GLM claim nor establishes responsive end-to-end coding performance.
2026-09-15T12:21:52Z
evidence attached: reddit.post.1wgy6tm โ Hands-on benchmark evidence supports the open local GLM/DeepSeek-style long-context coding inference hypothesis, including substantial decode and speculative-decoding gains.
2026-09-10T08:27:29Z
The refreshed comments repeat previously assessed comparisons, testing intentions and deployment questions; they add no independent reproduction or usable coding-agent result. The branch remains an author-reported optimization worth checking at a slower cadence, not yet evidence for changing local-backend choices.
2026-09-09T20:38:56Z
The new question about running a smaller quant on a 256GB machine highlights an unresolved deployment constraint, but supplies neither a compatibility result nor a reproduction. The case remains an author-reported optimization, with long-context memory requirements, coding quality and end-to-end latency still unestablished.
2026-09-09T17:26:08Z
The refreshed discussion adds a request for continuous batching and questions about output quality, not a reproduction or a new implementation result. This remains an author-reported local coding optimization with unresolved memory, latency and quality tradeoffs; the repository echo is not independent corroboration.
2026-09-09T15:37:58Z
The refreshed discussion adds a throughput comparison on four RTX 3090s, not a reproduction of the M3 Ultra branch or its long-context coding workload. The implementation remains worth tracking, but coding quality, memory requirements and end-to-end latency are still unestablished; interest in testing has not yet become validation.
2026-09-09T13:34:00Z
The implementation-linked claim remains worth watching, but this update adds no replication or evidence of usable end-to-end coding performance. The GitHub echo repeats the author's testimony rather than independently corroborating it, and a commenter intending to test is not yet a result.
2026-09-09T13:27:36Z
grounded: converges/low โ The reported hardware-specific Q4 implementation converges with Scottโs hardware-aware local-inference approach and is relevant to the local-backend option in h
2026-09-09T13:24:52Z
case created โ A repository-backed, workload-specific optimization claim is distinct from the existing GLM-5.3 model-capability case.