LocalLLaMA builder darklordfireape claims his MIT-licensed llama-halo-hybrid โ placing dense layers and KV on an R9700-class eGPU beside a Strix Halo APU โ now sustains >60 tok/s decode with 2000+ tok/s prefill at full 256K context, beating DGX Spark on cost-performance; independent replication and adoption would establish APU+eGPU hybrids as a replicated tier for long-context local inference.
state: corroboratedheat: lowuncertainty: mediumconvergesscott: highstrix-halo-local-inference hybrid-apu-gpu-inference local-inference-hardwaredarklordfireape
What is this?
The llama-halo-hybrid project (github.com/sixvolts/llama-halo-hybrid) is an MIT-licensed llama.cpp fork that implements a hybrid inference layout for AMD Strix Halo APUs paired with discrete Radeon GPUs (R9700-class eGPUs). The technique places dense layers and KV cache on the eGPU while keeping bulk model weights in the APU's unified memory, enabling 256K context windows. The project's README claims this single Strix Halo + R9700 combination beats an NVIDIA DGX Spark (GB10, 128GB, 273 GB/s) on cost-performance for Qwen3.8-Flash-Next, citing 2,000+ tok/s prefill and ~68 tok/s decode at 256K context. A second builder (Edenar on Reddit) has independently corroborated the core APU+eGPU hybrid architecture using a CMP 170HX eGPU, achieving ~100 tok/s aggregate on 27B models. The specific DGX Spark comparison and 256K-context numbers remain single-source (sixvolts repo).
Why it matters to Scott
The llama-halo-hybrid is a concrete, independently corroborated implementation of Scott's hardware-aware-local-inference concept โ placement-as-runtime-policy (dense layers + KV on eGPU, bulk weights in APU unified memory) with measured 256K-context results. A second builder (Edenar, CMP 170HX) replicates the core architecture on different hardware, moving the pattern from single-source claim to replicated tier. This directly validates the inference-field thesis (long context enlarges the operative world) and the sovereign-software-assurance vector (open-weight, local, vendor-independent). It bears on Scott's gamepc/Ollama substrate and opens a dated-receipts publishing window: the community is converging on the placement policy he argued for.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaip:concept.inference-fieldip:source.the-inference-field-ebookip:framework.code-first-architecturedev:concept.software-defined-heterogeneous-core-schedulingdev:project.linux-cpuradar:concept.local-inferenceradar:concept.inference-economicsradar:concept.llama-cppradar:concept.inference-optimizationradar:concept.long-context-inferenceradar:concept.agent-memoryradar:ex-cloud-gaming-gpu-local-inferenceradar:qwen38-flash-next-commodity-local-inferenceradar:llamacpp-fork-fragmentationradar:llamampere-long-context-throughput
queries asked of Scott's wikis
- hybrid placement runtime policy dense layers KV cache eGPU
- open-weights local inference hardware sovereignty cost-performance
- model sovereignty regulation local inference economics
- agent memory long-context local inference 256K
- hardware-aware inference placement-as-policy llama.cpp forks
Measured heat
now 0 pts/hpeak 7 pts/hcomments 0/hpeers p25momentum: steady2 platformsage 290h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p60 vs 1188 stories at the 168h mark (now 290h old) โ ahead of local-kv-cache-pressure-probe (1.0x), behind artificium-covering-design-search (1.0x)
Evidence (3) โ โญ canonical anchor
Interpretation history
2026-10-10T02:58:28Z
grounded: converges/high โ The llama-halo-hybrid is a concrete, independently corroborated implementation of Scott's hardware-aware-local-inference concept โ placement-as-runtime-policy (
2026-10-10T02:50:25Z
Independent user (Edenar) reports working Strix Halo + CMP 170HX eGPU hybrid achieving 100 tok/s aggregate on 27B models โ different hardware, same architecture โ corroborating the core APU+eGPU hybrid inference tier. Headline claims (>60 tok/s decode, 2000+ prefill at 256K, DGX Spark comparison) remain single-source.
2026-10-10T01:44:11Z
evidence attached: reddit.post.1x1z0mx โ User-reported Strix Halo + CMP 170HX eGPU hybrid setup provides real-world corroboration for the APU+eGPU hybrid inference tier.
2026-09-30T22:40:10Z
origin walked (opencode/cheap-glm, conf 0.9): anchor reddit.post.1wuet5g -> echo.github.53d1995b06 by sixvolts
2026-09-30T22:00:30Z
grounded: converges/medium โ Converges with his hardware-aware-local-inference position: llama-halo-hybrid is placement-as-runtime-policy in exactly his sense โ dense layers and KV cache de
2026-09-30T21:52:31Z
case created โ Working open artifact with concrete measured claims directly targeting DGX Spark cost-performance โ a distinct Strix Halo episode no open case covers.
Decision trace
- 10-10 15:28attention_communicatedLocalLLaMA builder Edenar independently replicated the core llama-halo-hybrid architecture โ dense layers + KV cache on eGPU (CMP 170HX via USB4/PCIe 2.0 x4), bulk weights in Strix Halo APU unified me
- 10-10 15:28attention_routePrevious judgment was watch (low confidence) pending replication. Edenar's independent build on different hardware constitutes the consequential change: the core hardware-aware placement policy S
- 10-10 14:01attention_routePrevious judgment was watch (low confidence) pending replication. Edenar's independent build on different hardware (CMP 170HX, USB4 eGPU dock) constitutes the consequential change: the core hardw
- 10-10 13:58attention_candidatematerial_reprice
- 10-10 13:58repriceIndependent user (Edenar) reports working Strix Halo + CMP 170HX eGPU hybrid achieving 100 tok/s aggregate on 27B models โ different hardware, same architecture โ corroborating the core APU+eGPU hybri
- 10-10 13:58groundThe llama-halo-hybrid is a concrete, independently corroborated implementation of Scott's hardware-aware-local-inference concept โ placement-as-runtime-policy (dense layers + KV on eGPU, bulk wei
- 10-10 12:47attention_routeThe editor compared this story and chose to keep watching.
- 10-10 12:44attention_candidateattach
- 10-10 12:44attachUser-reported Strix Halo + CMP 170HX eGPU hybrid setup provides real-world corroboration for the APU+eGPU hybrid inference tier.
- 10-10 12:38propose_attachUser-reported Strix Halo + CMP 170HX eGPU hybrid setup provides real-world corroboration for the APU+eGPU hybrid inference tier.
- 10-02 21:28review_screenjev screen: no material development (noul=0.19)
- 10-01 20:21sensor_dirtycomment_update
- 10-01 10:22sensor_dirtycomment_update
- 10-01 08:40promote_anchororigin walk conf 0.9
- 10-01 08:00groundConverges with his hardware-aware-local-inference position: llama-halo-hybrid is placement-as-runtime-policy in exactly his sense โ dense layers and KV cache deliberately parked on the bandwidth-rich
- 10-01 07:52createWorking open artifact with concrete measured claims directly targeting DGX Spark cost-performance โ a distinct Strix Halo episode no open case covers.