2026-10-11 16:38 UTC

Redditor deathcom65 reports that nasone32's specialized llama.cpp fork raises Qwen3.8 Q8 decode throughput from about 28 to 82 tokens per second at 60K context on dual Radeon 7900 XTX GPUs, potentially making long-context local agents substantially more responsive on consumer AMD hardware.

state: watchingheat: lowuncertainty: highknownscott: lowlocal-inference amd-gpu llama-cppdeathcom65nasone32

What is this?

A Redditor (deathcom65) reports that nasone32's RDNA3-tuned llama.cpp fork lifts Qwen3.8-27B Q8_0 decode from ~28 to ~82 tokens/s at 60K context on dual Radeon 7900 XTX cards; the supplied web results do not mention or verify that fork directly. What they do establish is that the presumed mechanism is now mainstream: llama.cpp has merged multi-token prediction, with paired community benchmarks showing ~1.7–2.4× decode gains on dense 27B models (including Q8_0 and AMD hardware), and an independent r/ROCm post reports 29→69 t/s on the same dual-7900-XTX rig class using tensor split + MTP — at 4-bit, with Q8 noted as ~7% slower and context-halving to fit. Around this, tuned-AMD results run higher still (a drafter-based stack claims 227 t/s on a single R9700; an MTP-flag recipe repo documents +33–39% with paired benchmarks and a known unsafe-math correctness footgun), and multiple independent AMD inference efforts (ZINC, ik_llama.cpp) now beat mainline llama.cpp. Net: the 82 t/s headline sits above every independent public result at its exact quant/context but inside a now-demonstrated envelope; the fork itself, its output correctness, and the Q8/60K figure remain unreplicated.

Why it matters to Scott

Scott's own wikis already carry this frame — dev:concept.hardware-aware-local-inference treats precision, placement and compilation as explicit runtime policy, and the new grounding (MTP merged into mainline llama.cpp with independent ~1.7–2.4× paired benchmarks, plus a same-rig 29→69 t/s replication) resolves the fork's mechanism question without touching anything he runs: gamepc is WSL2/CUDA/Ollama and ask's local path is secondary with native tools suppressed. The development is real but lands in radar lineage — fork-fragmentation and mainline adaptive-MTP questions — not in Scott's canon; the world again exemplifying his runtime-policy pattern is not news for him.
dev:concept.hardware-aware-local-inferencedev:project.gamepcradar:concept.speculative-decodingradar:concept.amd-inferenceradar:llamacpp-fork-fragmentationradar:llama-cpp-adaptive-mtp
queries asked of Scott's wikis
  • local inference stack Ollama CUDA gamepc positions
  • speculative decoding MTP acceptance quality tradeoffs
  • AMD ROCm consumer GPU inference notes
  • long-context local coding agent hosting feasibility
  • local model economics vs API serving positions
  • llama.cpp forks and tooling ecosystem notes

Measured heat

now 0 pts/hpeak 1 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 842h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-06 14:00⭐ origin echo-reconstructedThe linked repository is the primary artifact: “This is the llama.cpp build I use for Qwen on AMD RDNA3,” tested with “two RX 7900 XTX cards
nasone32 on github (echo) · attributed from reddit.post.1wjd5hf
—
09-18 02:00first on r/LocalLLaMA · published · +276.0hdual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode
deathcom65
—
09-18 02:00amplified on r/LocalLLaMA 👑reddit.post.1wjd5hf
deathcom65
peak 43 · 29 comments · 59% of case engagement
09-19 19:44amplified on r/LocalLLaMAreddit.post.1wkvl7h
W61k3r
peak 8 · 19 comments · 22% of case engagement
10-06 23:06amplified on r/LocalLLaMAreddit.post.1wzh32j
NoFee9147
peak 11 · 12 comments · 19% of case engagement
09-18 02:20our radar first saw it · +276.3hdiscovery anchor: reddit.post.1wjd5hf—
pace: p68 vs 519 stories at the 720h mark (now 842h old) — ahead of openai-german-wiki-incident (1.0x), behind openai-chatgpt-mil-genai-deployment (1.0x)

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditdual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode
LocalLLaMA
deathcom654329
🟧 echo.github ⭐The linked repository is the primary artifact: “This is the llama.cpp build I use for Qwen on AMD RDNA3,” tested with “two RX 7900 XTX cardsnasone32——
🟠 redditFinally got qwen 3.8 q4 k_M 27b runnin slick 40tkps load at 240k inf 4qnl for stable agentic coding on a 7900xtx. I have been finally able to give hermes a task and come back to results. Built a panel to manage my inference servers from hermes.
LocalLLaMA
W61k3r819
🟠 redditOptimizations Claude did for Qwen 3.8 27B and Qwen Flash Next on dual and quad 7900xtx
LocalLLaMA
NoFee91471112

Interpretation history

Decision trace