Redditor deathcom65 reports that nasone32's specialized llama.cpp fork raises Qwen3.8 Q8 decode throughput from about 28 to 82 tokens per second at 60K context on dual Radeon 7900 XTX GPUs, potentially making long-context local agents substantially more responsive on consumer AMD hardware.
state: watchingheat: lowuncertainty: highknownscott: lowlocal-inference amd-gpu llama-cppdeathcom65nasone32
What is this?
A Redditor (deathcom65) reports that nasone32's RDNA3-tuned llama.cpp fork lifts Qwen3.8-27B Q8_0 decode from ~28 to ~82 tokens/s at 60K context on dual Radeon 7900 XTX cards; the supplied web results do not mention or verify that fork directly. What they do establish is that the presumed mechanism is now mainstream: llama.cpp has merged multi-token prediction, with paired community benchmarks showing ~1.7–2.4× decode gains on dense 27B models (including Q8_0 and AMD hardware), and an independent r/ROCm post reports 29→69 t/s on the same dual-7900-XTX rig class using tensor split + MTP — at 4-bit, with Q8 noted as ~7% slower and context-halving to fit. Around this, tuned-AMD results run higher still (a drafter-based stack claims 227 t/s on a single R9700; an MTP-flag recipe repo documents +33–39% with paired benchmarks and a known unsafe-math correctness footgun), and multiple independent AMD inference efforts (ZINC, ik_llama.cpp) now beat mainline llama.cpp. Net: the 82 t/s headline sits above every independent public result at its exact quant/context but inside a now-demonstrated envelope; the fork itself, its output correctness, and the Q8/60K figure remain unreplicated.
Why it matters to Scott
Scott's own wikis already carry this frame — dev:concept.hardware-aware-local-inference treats precision, placement and compilation as explicit runtime policy, and the new grounding (MTP merged into mainline llama.cpp with independent ~1.7–2.4× paired benchmarks, plus a same-rig 29→69 t/s replication) resolves the fork's mechanism question without touching anything he runs: gamepc is WSL2/CUDA/Ollama and ask's local path is secondary with native tools suppressed. The development is real but lands in radar lineage — fork-fragmentation and mainline adaptive-MTP questions — not in Scott's canon; the world again exemplifying his runtime-policy pattern is not news for him.
dev:concept.hardware-aware-local-inferencedev:project.gamepcradar:concept.speculative-decodingradar:concept.amd-inferenceradar:llamacpp-fork-fragmentationradar:llama-cpp-adaptive-mtp
queries asked of Scott's wikis
- local inference stack Ollama CUDA gamepc positions
- speculative decoding MTP acceptance quality tradeoffs
- AMD ROCm consumer GPU inference notes
- long-context local coding agent hosting feasibility
- local model economics vs API serving positions
- llama.cpp forks and tooling ecosystem notes
Measured heat
now 0 pts/hpeak 1 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 842h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p68 vs 519 stories at the 720h mark (now 842h old) — ahead of openai-german-wiki-incident (1.0x), behind openai-chatgpt-mil-genai-deployment (1.0x)
Evidence (4) — ⭐ canonical anchor
Interpretation history
2026-10-07T00:41:14Z
grounded: known/low — Scott's own wikis already carry this frame — dev:concept.hardware-aware-local-inference treats precision, placement and compilation as explicit runtime policy,
2026-10-07T00:33:25Z
NoFee9147's Claude-written ROCm/llama-server fixes with measured decode gains (24.6→36.3 t/s Flash-Next) make this a multi-track scene rather than a single-source report: long-context Qwen coding on multi-7900XTX rigs is being actively optimized by several independent efforts — but none replicates the 82 t/s headline, which now reads as an unexplained outlier above what independent tracks achieve. Promote to watching pending the promised GitHub push or third-party replication of nasone32's fork; Scott relevance unchanged.
2026-10-06T23:36:36Z
evidence attached: reddit.post.1wzh32j — Claude-written ROCm/llama-server fixes on the same multi-7900XTX Qwen3.8 throughput episode with concrete measured gains, a second independent optimization track on the same rig class.
2026-09-19T20:22:39Z
The new report adds an adjacent AMD coding-agent deployment, not a replication of nasone32’s dual-7900-XTX fork: it uses Q4 quantization on one Radeon with vision offloaded to an RTX 2060. This modestly supports workload feasibility but does not establish the headline speedup or justify changing Scott’s stack.
2026-09-19T20:22:14Z
evidence attached: reddit.post.1wkvl7h — Independent AMD-user evidence supports the case that tuned multi-GPU llama.cpp deployments can make long-context Qwen coding inference practical on consumer hardware.
2026-09-18T12:37:58Z
Discussion adds possible explanations and a pointer to earlier prompt-processing and memory improvements, but no controlled replication of the headline decode gain. Conflicting model-size assumptions make the bandwidth arguments inconclusive; this remains a narrow hardware-specific lead rather than evidence of better local-agent performance.
2026-09-18T02:27:24Z
grounded: known/low — This is another example of Scott’s existing “Hardware-aware local inference” approach, not a demonstrated extension: his documented gamepc stack uses CUDA/Ollam
2026-09-18T02:24:19Z
origin walked (codex/luna, conf 0.98): anchor reddit.post.1wjd5hf -> echo.github.2d2e3242df by nasone32
2026-09-18T02:22:17Z
case created — A public hardware-specific fork and quantified user measurement establish a distinct, reproducible optimization claim, although benchmark controls and correctness remain unverified.
Decision trace
- 10-11 14:29review_screenThe only change is a new comment expressing an opinion about llama.cpp maintainers and mentioning a switch to Strata. This adds no new factual evidence, replication data, technical details, or verific
- 10-11 14:29review_screenjev screen borderline (noul=0.49) — luna review
- 10-07 22:20sensor_dirtycomment_update
- 10-07 11:41repriceNoFee9147's Claude-written ROCm/llama-server fixes with measured decode gains (24.6→36.3 t/s Flash-Next) make this a multi-track scene rather than a single-source report: long-context Qwen coding
- 10-07 11:41groundScott's own wikis already carry this frame — dev:concept.hardware-aware-local-inference treats precision, placement and compilation as explicit runtime policy, and the new grounding (MTP merged i
- 10-07 10:36attachClaude-written ROCm/llama-server fixes on the same multi-7900XTX Qwen3.8 throughput episode with concrete measured gains, a second independent optimization track on the same rig class.
- 10-07 10:26propose_attachClaude-written ROCm/llama-server fixes on the same multi-7900XTX Qwen3.8 throughput episode with concrete measured gains, a second independent optimization track on the same rig class.
- 09-25 19:23review_screenjev screen: no material development (noul=0.16)
- 09-22 08:43review_screenjev screen: no material development (noul=0.26)
- 09-20 06:22repriceThe new report adds an adjacent AMD coding-agent deployment, not a replication of nasone32’s dual-7900-XTX fork: it uses Q4 quantization on one Radeon with vision offloaded to an RTX 2060. This modest
- 09-20 06:22attachIndependent AMD-user evidence supports the case that tuned multi-GPU llama.cpp deployments can make long-context Qwen coding inference practical on consumer hardware.
- 09-20 06:21propose_attachIndependent AMD-user evidence supports the case that tuned multi-GPU llama.cpp deployments can make long-context Qwen coding inference practical on consumer hardware.
- 09-19 08:30review_screenThe added author comment only describes a planned EXL3 port and aspirational 48GB target; it provides no released implementation, benchmark, or independent corroboration of the headline speedup.
- 09-19 08:20sensor_dirtycomment_update
- 09-18 22:37repriceDiscussion adds possible explanations and a pointer to earlier prompt-processing and memory improvements, but no controlled replication of the headline decode gain. Conflicting model-size assumptions
- 09-18 22:37review_screenThe new comments add a possible explanation and reference an earlier report of prompt-speed and OOM improvements, but they provide no controlled benchmark or clear first-hand verification beyond the e
- 09-18 21:21sensor_dirtycomment_update
- 09-18 16:29review_screenThe new comments only express skepticism and repeat existing concerns about benchmark plausibility, controls, and repository trust; they add no verified result or consequential new fact.
- 09-18 14:21sensor_dirtycomment_update
- 09-18 12:27groundThis is another example of Scott’s existing “Hardware-aware local inference” approach, not a demonstrated extension: his documented gamepc stack uses CUDA/Ollama, and the hits establish no dual-Radeon
- 09-18 12:24promote_anchororigin walk conf 0.98
- 09-18 12:22createA public hardware-specific fork and quantified user measurement establish a distinct, reproducible optimization claim, although benchmark controls and correctness remain unverified.