2026-10-11 17:12 UTC

whodoneit1 claims their released vLLM modifications convert NVFP4 weights online to an MXFP4 fast path and run Qwen3.8 27B on AMD R9700 hardware at 5,809 prefill and 276 decode tokens per second, potentially improving practical AMD local-inference throughput.

state: expiredheat: lowuncertainty: highknownscott: lowlocal-inference amd-gpu quantizationwhodoneit1GGZ14
Surfaced 2026-09-21T15:45:27Z — Optimizing Qwen3.8 27B on one AMD R9700 — An independent dual-R9700 hands-on setup adds practical evidence that the broader ROCm/vLLM path is usable, but it does not reproduce the online NVFP4-to-MXFP4 conversion or headline throughput. The loud multi-platform reading raises attention to medium, not high, because spread remains concentrated around the original claims and one related configuration rather than broad exact-path reproductions.

What is this?

The case concerns community modifications to vLLM that whodoneit1 claims convert NVFP4 model weights online to an MXFP4 fast path, achieving 5,809 prefill and 276 decode tokens per second for Qwen3.8 27B on AMD R9700 hardware. A supplied blog snippet describes a related patched vLLM “Radiance” build using MXFP4 and speculative decoding on dual R9700s; its author reports 3,948 prefill and 184.9 generation tokens per second, while citing a Reddit report of roughly 280. This supports interest in the optimization path, but the snippets do not establish the claimed online conversion implementation, reproduce the exact headline results, or verify whodoneit1’s authorship or GGZ14’s role.

Why it matters to Scott

Scott’s “Hardware-aware local inference” page already treats numerical precision and accelerator-specific execution as runtime policy; this is another example of that position, not an established extension or challenge to it. His documented gamepc stack uses CUDA/Ollama, with no supplied evidence of an AMD/vLLM deployment or migration plan, and the conversion and headline throughput remain unverified; the radar’s “vllm-amd-speculative-decoding” page tracks a related optimization, not this exact development.
dev:concept.hardware-aware-local-inferencedev:project.gamepcradar:vllm-amd-speculative-decodingradar:concept.amd-inferenceradar:concept.quantization
queries asked of Scott's wikis
  • Local inference hardware economics and AMD GPU deployment
  • vLLM serving stack custom kernels and maintenance costs
  • Coding agent prefill latency versus decode throughput requirements
  • Low-bit quantization quality tradeoffs and model portability
  • Speculative decoding benchmarks and realistic agent workloads

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

09-17 11:50⭐ origin directly observedOptimizing Qwen3.8 27B on one AMD R9700
Eliovp on hacker news
—
09-15 20:51first on r/LocalLLaMA · published · +-39.0hNVFP4 Qwen3.8 27b - AMD R9700's running @ 5,809 tok/s Prefill and 276 tok/s decode. Leveraging my MXFP4 fast path.
whodoneit1
—
09-15 21:47first on github (echo) · first seen by us · +-38.1hThe author links this implementation while reporting newly added online NVFP4-to-MXFP4 conversion and the stated Qwen3.8 27B throughput on R
GGZ14 / whodoneit1
—
09-15 20:51amplified on r/LocalLLaMAreddit.post.1whcik4
whodoneit1
peak 8 · 90 comments · 19% of case engagement
09-17 11:50amplified on hacker newshn.story.49739432
Eliovp
peak 1 · 0 comments · 0% of case engagement
09-17 15:14amplified on r/LocalLLaMA 👑reddit.post.1wiws8e
whodoneit1
peak 200 · 112 comments · 62% of case engagement
09-21 00:57amplified on r/LocalLLaMAreddit.post.1wlyc2p
Pyrolistical
peak 10 · 16 comments · 5% of case engagement
09-23 17:49amplified on r/LocalLLaMAreddit.post.1wod07t
neuromacmd
peak 4 · 8 comments · 2% of case engagement
09-24 14:02amplified on r/LocalLLaMAreddit.post.1wp2hqk
Public_Umpire_1099
peak 47 · 10 comments · 11% of case engagement
09-15 21:20our radar first saw it · +-38.5hdiscovery anchor: reddit.post.1whcik4—
09-21 15:45reached heat=high · +99.9h · via ledger——

Evidence (7) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditNVFP4 Qwen3.8 27b - AMD R9700's running @ 5,809 tok/s Prefill and 276 tok/s decode. Leveraging my MXFP4 fast path.
LocalLLaMA
whodoneit1890
🟧 echo.githubThe author links this implementation while reporting newly added online NVFP4-to-MXFP4 conversion and the stated Qwen3.8 27B throughput on RGGZ14 / whodoneit1——
🟧 hn ⭐Optimizing Qwen3.8 27B on one AMD R9700Eliovp10
🟠 reddit153 tok/s on 1x AMD Radeon R9700 running Qwen3.8 27b NVFP4, 470 tok/s @ 8 conc requests, Prefill @ 3,619 tok/s
LocalLLaMA
Retrieved article excerpt

Open article · Retrieved 2026-09-17T15:21:50.578923+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. © "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
whodoneit1197112
🟠 reddit2x r9700 in one strix halo machine
LocalLLaMA
Pyrolistical1116
🟠 redditDeepSeek-V4-Flash-0731 at ~40–50 tok/s on 2× Radeon AI PRO R9700 with the affinity engine (prebuilt quant + fixes)
LocalLLaMA
neuromacmd1111
🟠 redditR9V Update: Created and adopted KVA projections based on Deepseek V4.1 Flash + HySparse2/MiMo-V3 for Qwen3.8 Flash Next. This is a game changer for models that don't natively implement it. 1.45-1.85x speedup in prefill to 3k+ at a small deficit to perplexity. [2x R9700, 128GB DDR5]
LocalLLaMA
Public_Umpire_10994710

Interpretation history

Decision trace