2026-10-11 16:38 UTC

LocalLLaMA builder darklordfireape claims his MIT-licensed llama-halo-hybrid โ€” placing dense layers and KV on an R9700-class eGPU beside a Strix Halo APU โ€” now sustains >60 tok/s decode with 2000+ tok/s prefill at full 256K context, beating DGX Spark on cost-performance; independent replication and adoption would establish APU+eGPU hybrids as a replicated tier for long-context local inference.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: highstrix-halo-local-inference hybrid-apu-gpu-inference local-inference-hardwaredarklordfireape

What is this?

The llama-halo-hybrid project (github.com/sixvolts/llama-halo-hybrid) is an MIT-licensed llama.cpp fork that implements a hybrid inference layout for AMD Strix Halo APUs paired with discrete Radeon GPUs (R9700-class eGPUs). The technique places dense layers and KV cache on the eGPU while keeping bulk model weights in the APU's unified memory, enabling 256K context windows. The project's README claims this single Strix Halo + R9700 combination beats an NVIDIA DGX Spark (GB10, 128GB, 273 GB/s) on cost-performance for Qwen3.8-Flash-Next, citing 2,000+ tok/s prefill and ~68 tok/s decode at 256K context. A second builder (Edenar on Reddit) has independently corroborated the core APU+eGPU hybrid architecture using a CMP 170HX eGPU, achieving ~100 tok/s aggregate on 27B models. The specific DGX Spark comparison and 256K-context numbers remain single-source (sixvolts repo).

Why it matters to Scott

The llama-halo-hybrid is a concrete, independently corroborated implementation of Scott's hardware-aware-local-inference concept โ€” placement-as-runtime-policy (dense layers + KV on eGPU, bulk weights in APU unified memory) with measured 256K-context results. A second builder (Edenar, CMP 170HX) replicates the core architecture on different hardware, moving the pattern from single-source claim to replicated tier. This directly validates the inference-field thesis (long context enlarges the operative world) and the sovereign-software-assurance vector (open-weight, local, vendor-independent). It bears on Scott's gamepc/Ollama substrate and opens a dated-receipts publishing window: the community is converging on the placement policy he argued for.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaip:concept.inference-fieldip:source.the-inference-field-ebookip:framework.code-first-architecturedev:concept.software-defined-heterogeneous-core-schedulingdev:project.linux-cpuradar:concept.local-inferenceradar:concept.inference-economicsradar:concept.llama-cppradar:concept.inference-optimizationradar:concept.long-context-inferenceradar:concept.agent-memoryradar:ex-cloud-gaming-gpu-local-inferenceradar:qwen38-flash-next-commodity-local-inferenceradar:llamacpp-fork-fragmentationradar:llamampere-long-context-throughput
queries asked of Scott's wikis
  • hybrid placement runtime policy dense layers KV cache eGPU
  • open-weights local inference hardware sovereignty cost-performance
  • model sovereignty regulation local inference economics
  • agent memory long-context local inference 256K
  • hardware-aware inference placement-as-policy llama.cpp forks

Measured heat

now 0 pts/hpeak 7 pts/hcomments 0/hpeers p25momentum: steady2 platformsage 290h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-29 14:00โญ origin echo-reconstructedREADME of the llama-halo-hybrid repo (a llama.cpp fork): "This is a fork of llama.cpp that builds out support for Strix Halo with a GPU side
sixvolts on github (echo) ยท attributed from reddit.post.1wuet5g
โ€”
09-30 19:46first on r/LocalLLaMA ยท published ยท +29.8hUpdate: Strix Halo + R9700 with llama-halo-hybrid - now beats DGX Spark
darklordfireape
โ€”
09-30 19:46amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wuet5g
darklordfireape
peak 24 ยท 28 comments ยท 80% of case engagement
10-09 22:32amplified on r/LocalLLaMAreddit.post.1x1z0mx
Edenar
peak 6 ยท 7 comments ยท 20% of case engagement
09-30 21:20our radar first saw it ยท +31.3hdiscovery anchor: reddit.post.1wuet5gโ€”
pace: p60 vs 1188 stories at the 168h mark (now 290h old) โ€” ahead of local-kv-cache-pressure-probe (1.0x), behind artificium-covering-design-search (1.0x)

Evidence (3) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditUpdate: Strix Halo + R9700 with llama-halo-hybrid - now beats DGX Spark
LocalLLaMA
darklordfireape2427
๐ŸŸง echo.github โญREADME of the llama-halo-hybrid repo (a llama.cpp fork): "This is a fork of llama.cpp that builds out support for Strix Halo with a GPU sidesixvoltsโ€”โ€”
๐ŸŸ  redditStrix halo + CMP 170HX setup
LocalLLaMA
Edenar67

Interpretation history

Decision trace