2026-10-11 16:37 UTC

strix-halo

band: warmmomentum: stable score: 0.576
temperature history

Episodes (3)

Redditor einthecorgi2 reports that the released Atlas inference engine works locally and supports Strix Halo, potentially providing an alternative to llama.cpp on that hardware.
resolvednovelscott: low
LocalLLaMA builder Yaniss916 claims the released Kyojin ROCm engine (built on ExLlamaV3 for gfx1151) and mixed EXL3 packs run 300B-class MoE models โ€” GLM-5.3-Flash at ~580 tok/s prefill and 26โ€“30 tok/s decode, MiMo-V2.6-Flash up to 44 tok/s โ€” with near-FP8 fidelity (KLD 0.151/0.071, ~90% top-1 agreement) on a single 128GB Strix Halo mini PC; independent replication and adoption would establish Strix Halo APUs as a standard tier for large-MoE local inference.
corroboratedconvergesscott: high
stereohype claims Halogen 0.16+'s OpenAI-compatible endpoints over Strix Halo's idle XDNA2 NPU โ€” with measured 15/20-vs-9/20 semantic search over grep at 70โ€“130ms on 0.17.1 and a claimed 30x NPU latency cut โ€” make the NPU a working auxiliary small-model tier (search, dedup, decisions, injection screening) inside coding agents; adoption of NPU-backed components in other local agent stacks confirms it, quiet fade closes it.
seedconvergesscott: high