2026-10-11 16:38 UTC

LocalLLaMA builder Yaniss916 claims the released Kyojin ROCm engine (built on ExLlamaV3 for gfx1151) and mixed EXL3 packs run 300B-class MoE models β€” GLM-5.3-Flash at ~580 tok/s prefill and 26–30 tok/s decode, MiMo-V2.6-Flash up to 44 tok/s β€” with near-FP8 fidelity (KLD 0.151/0.071, ~90% top-1 agreement) on a single 128GB Strix Halo mini PC; independent replication and adoption would establish Strix Halo APUs as a standard tier for large-MoE local inference.

state: corroboratedheat: mediumuncertainty: mediumconvergesscott: highlocal-inference strix-halo moe-quantizationYaniss916

What is this?

Yamz Labs (Reddit u/Yaniss916) released Kyojin, an open-source ROCm inference engine built on ExLlamaV3 for AMD's Strix Halo APU (gfx1151), claiming it runs two ~300B-class MoE models β€” GLM-5.3-Flash and MiMo-V2.6-Flash β€” on a single 128GB mini PC with published throughput numbers (GLM: ~580 tok/s prefill, 26–30 tok/s decode; MiMo: up to 44 tok/s decode with speculative decoding) and near-FP8 fidelity metrics (KLD 0.151/0.071, ~90% top-1 agreement). Mixed EXL3 quantization packs and engine source are public on GitHub and Hugging Face. The hardware-tier claim β€” that Strix Halo APUs constitute a practical standard tier for large-MoE local inference β€” is supported by the builder's own measurements and subsequent independent user reports on sibling engines (Halogen, Gufo/Strixite, Radiance), but Kyojin's specific headline numbers remain builder-measured only; a git-history provenance challenge (alleged rewrite of vcruz305/exllamav3-amd) is unanswered, and no independent numeric replication of the 580/26–44 tok/s claims exists.

Why it matters to Scott

Converges with Scott's hardware-aware local inference doctrine (dev:concept.hardware-aware-local-inference): a new consumer-APU tier (Strix Halo, 128GB unified memory) now has multiple independent engines (Kyojin, Halogen, Gufo/Strixite, Radiance) that treat the accelerator's characteristics β€” ROCm path, mixed EXL3 quantization, MoE expert routing β€” as explicit runtime policy, exactly the pattern the doctrine predicts. The tier economics shift (300B-class MoE on a mini PC) bears on his gamepc substrate (dev:project.gamepc) and single-tenant appliance targets (dev:concept.single-tenant-ai-appliance, dev:project.openclaw).
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:concept.single-tenant-ai-appliancedev:project.openclawradar:atlas-local-inference-engineradar:amd-llama-cpp-prefill-speedupradar:adaptive-kv-cache-streamingradar:afm3-prompt-conditioned-pruning
queries asked of Scott's wikis
  • hardware-aware local inference doctrine Strix Halo AMD APU
  • MoE quantization EXL3 mixed-precision local inference economics
  • model sovereignty open weights local-first AI hardware tiers
  • ExLlamaV3 turboderp engine lineage ROCm gfx1151
  • local inference hardware tier strategy consumer APU vs GPU
  • RAG knowledge systems local model serving Strix Halo

Measured heat

now 31 pts/hpeak 166 pts/hcomments 9/hpeers p95momentum: cooling3 platformsage 242h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

10-01 14:00⭐ origin echo-reconstructedREADME: "Kyojin is the Yamz inference engine for AMD Strix Halo, built on ExLlamaV3 by turboderp. It adds a ROCm decode and prefill path for
Yamz Labs (Reddit u/Yaniss916) on github (echo) Β· attributed from reddit.post.1wwocik
β€”
10-03 14:16first on r/LocalLLaMA Β· published Β· +48.3hTwo ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine
Yaniss916
β€”
10-11 01:43first on hacker news Β· published Β· +227.7hGMKtec EVO-X3 Mini-PC Review Looking at a 128GB AMD Local AI Box
lumpa
β€”
10-03 14:16amplified on r/LocalLLaMAreddit.post.1wwocik
Yaniss916
peak 65 Β· 58 comments Β· 8% of case engagement
10-05 15:25amplified on r/LocalLLaMAreddit.post.1wybesy
Yaniss916
peak 111 Β· 59 comments Β· 12% of case engagement
10-08 11:07amplified on r/LocalLLaMAreddit.post.1x0o2yy
deepu105
peak 33 Β· 76 comments Β· 7% of case engagement
10-10 16:25amplified on r/LocalLLaMAreddit.post.1x2jf0y
forevergeeks
peak 0 Β· 7 comments Β· 0% of case engagement
10-11 00:20amplified on r/LocalLLaMA πŸ‘‘reddit.post.1x2u742
sayamss
peak 732 Β· 312 comments Β· 71% of case engagement
10-11 01:43amplified on hacker newshn.story.50039003
lumpa
peak 4 Β· 0 comments Β· 0% of case engagement
10-03 14:20our radar first saw it Β· +48.3hdiscovery anchor: reddit.post.1wwocikβ€”
pace: p79 vs 1188 stories at the 168h mark (now 242h old) β€” ahead of pennsylvania-datacenter-opposition (1.0x), behind llama-cpp-specdec-moe-fusion (1.0x)

Evidence (7) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditTwo ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine
LocalLLaMA
Retrieved article excerpt

Open article Β· Retrieved 2026-10-03T14:24:42.030203+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. Β© "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
Yaniss9166458
🟧 echo.github ⭐README: "Kyojin is the Yamz inference engine for AMD Strix Halo, built on ExLlamaV3 by turboderp. It adds a ROCm decode and prefill path forYamz Labs (Reddit u/Yaniss916)β€”β€”
🟠 redditQwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open
LocalLLaMA
Yaniss91611159
🟠 redditHalogen + Qwen Flash Next keeps getting better
LocalLLaMA
deepu1053376
🟠 redditqwen3.6-35b-a3b on the Strix Halo with 64GB of memory
LocalLLaMA
forevergeeks07
🟠 redditBuilding a 4x R9700 setup for a 10 person startup
LocalLLaMA
sayamss732312
🟧 hnGMKtec EVO-X3 Mini-PC Review Looking at a 128GB AMD Local AI Boxlumpa40

Interpretation history

Decision trace