Yamz Labs (Reddit u/Yaniss916) released Kyojin, an open-source ROCm inference engine built on ExLlamaV3 for AMD's Strix Halo APU (gfx1151), claiming it runs two ~300B-class MoE models β GLM-5.3-Flash and MiMo-V2.6-Flash β on a single 128GB mini PC with published throughput numbers (GLM: ~580 tok/s prefill, 26β30 tok/s decode; MiMo: up to 44 tok/s decode with speculative decoding) and near-FP8 fidelity metrics (KLD 0.151/0.071, ~90% top-1 agreement). Mixed EXL3 quantization packs and engine source are public on GitHub and Hugging Face. The hardware-tier claim β that Strix Halo APUs constitute a practical standard tier for large-MoE local inference β is supported by the builder's own measurements and subsequent independent user reports on sibling engines (Halogen, Gufo/Strixite, Radiance), but Kyojin's specific headline numbers remain builder-measured only; a git-history provenance challenge (alleged rewrite of vcruz305/exllamav3-amd) is unanswered, and no independent numeric replication of the 580/26β44 tok/s claims exists.
Converges with Scott's hardware-aware local inference doctrine (dev:concept.hardware-aware-local-inference): a new consumer-APU tier (Strix Halo, 128GB unified memory) now has multiple independent engines (Kyojin, Halogen, Gufo/Strixite, Radiance) that treat the accelerator's characteristics β ROCm path, mixed EXL3 quantization, MoE expert routing β as explicit runtime policy, exactly the pattern the doctrine predicts. The tier economics shift (300B-class MoE on a mini PC) bears on his gamepc substrate (dev:project.gamepc) and single-tenant appliance targets (dev:concept.single-tenant-ai-appliance, dev:project.openclaw).
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:concept.single-tenant-ai-appliancedev:project.openclawradar:atlas-local-inference-engineradar:amd-llama-cpp-prefill-speedupradar:adaptive-kv-cache-streamingradar:afm3-prompt-conditioned-pruning
queries asked of Scott's wikis
- hardware-aware local inference doctrine Strix Halo AMD APU
- MoE quantization EXL3 mixed-precision local inference economics
- model sovereignty open weights local-first AI hardware tiers
- ExLlamaV3 turboderp engine lineage ROCm gfx1151
- local inference hardware tier strategy consumer APU vs GPU
- RAG knowledge systems local model serving Strix Halo
now 31 pts/hpeak 166 pts/hcomments 9/hpeers p95momentum: cooling3 platformsage 242h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
2026-10-11T02:46:03Z
grounded: converges/high β Converges with Scott's hardware-aware local inference doctrine (dev:concept.hardware-aware-local-inference): a new consumer-APU tier (Strix Halo, 128GB unified
2026-10-11T02:36:21Z
The Strix Halo hardware-tier hypothesis is now corroborated by four independent engines (Kyojin, Halogen, Gufo/Strixite family, Radiance fork), multiple user reports across memory tiers (64GBβ128GB), a production 4ΓR9700 deployment serving 16 concurrent sessions, and a hardware review. Kyojin's specific headline numbers (580 tok/s prefill, 26β44 tok/s decode) and the git-history provenance challenge remain single-poster and unanswered. The case has crossed from 'watching' to 'corroborated' on the tier hypothesis; the engine-specific claims are still gated.
2026-10-11T02:33:19Z
evidence attached: hn.story.50039003 β Hardware review of 128GB AMD Strix Halo mini-PC corroborates the platform's emergence as a practical tier for large-MoE local inference.
2026-10-11T01:03:20Z
New 4ΓR9700 (Strix Halo) deployment with published throughput numbers (Radiance fork, Qwen 3.8 Next Flash/27B + DSV4 Flash) adds a third independent engine corroborating the hardware-tier hypothesis. Kyojin's specific headline numbers (580 tok/s prefill, 26β44 tok/s decode) and the git-history provenance challenge remain unreplicated/unanswered. Engagement on Kyojin threads fully cooled; measured heat shows high historical percentile (89) but low current rate (1.8 pts/h).
2026-10-11T00:36:53Z
evidence attached: reddit.post.1x2u742 β Real-world 4ΓR9700 (Strix Halo) deployment running Qwen 3.8 Next Flash/27B and DSV4 Flash with published throughput numbers; corroborates Strix Halo as a practical local-inference tier.
2026-10-10T20:10:33Z
Second independent user report (forevergeeks: Qwen 3.6-35B-A3B at ~80 tok/s on 64GB Strix Halo) corroborates the hardware-tier hypothesis β Strix Halo as a standard tier for large-MoE local inference β but still no independent numeric replication of Kyojin's headline numbers (580 tok/s prefill, 26β44 tok/s decode). The git-history provenance challenge on Kyojin remains unanswered; Lemonade's ROCm-vs-Vulkan finding and Gufo precedent still weigh against single-poster claims. Sibling-engine field (halogen, gufo, strixite, ciru, strix-llama) active with Halogen thread growing discussion (56β76 comments) but no comparative benchmarks. Engagement on Kyojin threads fully cooled.
2026-10-10T18:47:50Z
evidence attached: reddit.post.1x2jf0y β User corroboration: Qwen 3.6-35B-A3B runs at ~80 tok/s on 64GB Strix Halo mini PC, supporting the case's claim about Strix Halo as a tier for large-MoE local inference.
2026-10-08T14:01:13Z
New independent user report (deepu105) confirms Strix Halo 128GB runs Qwen 3.8 Flash Next (~177B MoE) at ~45 tok/s via Halogen engine β corroborating the hardware-tier hypothesis (Strix Halo as standard tier for large MoE) but not the specific Kyojin engine claims. The git-history provenance challenge on Kyojin remains unanswered; Lemonade's ROCm-vs-Vulkan finding and Gufo precedent still weigh against single-poster numbers. Sibling-engine field (halogen, gufo, strixite, ciru, strix-llama) active but no comparative benchmarks. Engagement fully cooled (1.5 pts/h, 50th percentile, steady momentum).
2026-10-08T12:41:57Z
evidence attached: reddit.post.1x0o2yy β User report corroborates Strix Halo (128GB) running Qwen 3.8 Flash Next (~177B MoE) at 45 tps via Halogen, supporting the case's hypothesis about Strix Halo as a standard tier for large-MoE local inference.
2026-10-08T02:41:37Z
Engagement has decayed to near-zero velocity (0 pts/h, 17th-percentile peer rate) six days in; the brief velocity spike on the Qwen3.8-Flash-Next thread has fully cooled. Still only one community (Reddit) plus the builder's GitHub β no independent numeric replication of the 580/26β44 tok/s headlines, the git-history provenance challenge remains unanswered, and the crowded sibling-engine field (halogen, gufo, strixite, ciru, strix-llama) has no comparative benchmarks. The sustained same-team release cadence is real but does not yet constitute corroboration.
2026-10-05T18:57:47Z
The same team's second release (Qwen3.8-Flash-Next 125B, 44-59 tok/s, ~1,400 tok/s prefill, weights out, engine open) turns the case from one bold claim into a sustained release line on the tier β structurally stronger for the Strix-Halo-as-standard-tier hypothesis, yet still entirely first-party: no independent benchmark, the git-history provenance challenge stands unanswered, and the new thread itself reveals a crowded sibling-engine field (halogen, gufo, strixite, ciru, strix-llama) whose comparative numbers will be the real test. Heat to medium on a fresh release thread moving at 90th-percentile cohort rate, not high: no second community, no cross-platform spread, no independent implementation.
2026-10-05T17:32:20Z
evidence attached: reddit.post.1wybesy β Same team extends the Kyojin/Strix Halo line to Qwen3.8-Flash-Next 125B with strong published throughput and released weights β direct evidence for the Strix-Halo-as-standard-tier hypothesis.
2026-10-04T15:50:09Z
First independent hands-on signal: a stranger (ionizing) is running the MiMo EXL3 pack on this tier, reports it works with decent tool calls and feels faster than expected β but no benchmarks, so the 580/26β44 headline numbers remain the builder's own. A thread challenge that the repo is a git-history rewrite rather than an honest fork of vcruz305/exllamav3-amd adds an open provenance question; seedβwatching, and the brief 3.5x velocity spike has already decayed into a single engaged thread.
2026-10-03T14:47:49Z
origin walked (opencode/cheap-glm, conf 0.85): anchor reddit.post.1wwocik -> echo.github.61dded18c5 by Yamz Labs (Reddit u/Yaniss916)
2026-10-03T14:35:25Z
grounded: converges/medium β Converges with Scott's hardware-aware local-inference doctrine (dev:concept.hardware-aware-local-inference): a first-party gfx1151 engine that treats accelerato
2026-10-03T14:26:01Z
case created β First-party engine release with concrete KLD/throughput measurements on a Strix Halo tier no open case covers, and the builder's own post is the episode's origin.