2026-10-11 17:13 UTC

Extension-Bid-639 claims a build combining quantization, expert caching, host-RAM offload, and multi-token prediction raises full-261K-context Qwen3.8-Flash-Next decode throughput from 25–29 to 37–41 tokens per second on two RTX 3090 GPUs, potentially making long-context local coding inference practical on commodity multi-GPU systems.

state: resolvedheat: lowuncertainty: lowknownscott: mediumlocal-inference sparse-models speculative-decoding inference-economicsExtension-Bid-639llama.cppQwen

What is this?

A pseudonymous user, Extension-Bid-639, reports a build for Qwen3.8-Flash-Next that combines 4-bit quantization, expert caching, host-RAM offload, and multi-token prediction to raise full-261K-context decode speed from 25–29 to 37–41 tokens/second on two RTX 3090 GPUs. The Qwen model is described as a sparse 125B-parameter system with roughly 6B parameters active per token, making memory movement and PCIe topology important to performance. The supplied results support the architecture’s local-inference potential but do not directly verify this particular build, branch, hardware setup, or throughput claim; one source says independent verification of the new model’s performance is essentially nonexistent.

Why it matters to Scott

The radar already tracks this development through `radar:llama-cpp-hot-expert-offload` and `radar:llama-cpp-adaptive-mtp`; this update combines those techniques and claims a higher long-context throughput result rather than introducing a new direction. It still bears directly on Scott’s hardware-aware runtime policy and self-hosted GPU model zoo because the released branch may be testable, but the reported gain remains unverified.
dev:concept.hardware-aware-local-inferencedev:project.gamepcradar:llama-cpp-hot-expert-offloadradar:llama-cpp-adaptive-mtpradar:concept.long-context-inferenceradar:concept.moe-inference
queries asked of Scott's wikis
  • heterogeneous VRAM and host-RAM inference architecture
  • sparse MoE local inference economics
  • multi-token prediction versus speculative decoding
  • commodity multi-GPU coding-agent inference
  • long-context local inference practicality
  • memory bandwidth versus compute for local models

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

09-04 00:21⭐ origin directly observedUPDATE: Qwen3.8-Flash-Next on 2x3090 + DDR4 (Part 2): 25-29 -> 37-41 t/s decode (UD-Q4_K_XL + expert cache + MTP), plus a branch you can build
Extension-Bid-639 on r/LocalLLaMA
—
09-04 08:06first on r/LocalLLaMA · published · +7.7hQwen3.8-Flash-Next: 256k context, 16tok/s on DDR4 and a Tesla T4
BusTiny207
—
09-08 14:49first on hacker news · published · +110.5hBenchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
stared
—
09-08 19:17first on r/singularity · published · +114.9hQwen 3.8 27b with PI agent - pushed to its 3D graphic game limits - locally
Healthy-Nebula-3603
—
09-04 00:21amplified on r/LocalLLaMAreddit.post.1w6ozbj
Extension-Bid-639
peak 59 · 35 comments · 2% of case engagement
09-04 08:06amplified on r/LocalLLaMAreddit.post.1w6y38l
BusTiny207
peak 11 · 11 comments · 0% of case engagement
09-04 10:38amplified on r/LocalLLaMAreddit.post.1w70qrd
AppealSame4367
peak 3 · 23 comments · 0% of case engagement
09-04 12:40amplified on r/LocalLLaMAreddit.post.1w73aak
zRevengee
peak 64 · 27 comments · 2% of case engagement
09-04 16:17amplified on r/LocalLLaMAreddit.post.1w78wx4
Nota_ReAlperson
peak 2 · 25 comments · 1% of case engagement
09-04 16:28amplified on r/LocalLLaMAreddit.post.1w797w1
Forward_Jackfruit813
peak 4 · 9 comments · 0% of case engagement
53 more amplifiers in ainews.case_chain
09-04 01:20our radar first saw it · +1.0hdiscovery anchor: reddit.post.1w6ozbj—

Evidence (59) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐UPDATE: Qwen3.8-Flash-Next on 2x3090 + DDR4 (Part 2): 25-29 -> 37-41 t/s decode (UD-Q4_K_XL + expert cache + MTP), plus a branch you can build
LocalLLaMA
Extension-Bid-6396335
🟠 redditQwen3.8-Flash-Next: 256k context, 16tok/s on DDR4 and a Tesla T4
LocalLLaMA
BusTiny2071111
🟠 redditFastest qwen3.8 Flash Next Setup?
LocalLLaMA
AppealSame4367323
🟠 redditQwen 3.8 Flash Next Can Build Funny Games
LocalLLaMA
zRevengee5827
🟠 redditQwen 3.8 slow?
LocalLLaMA
Nota_ReAlperson025
🟠 redditQwen 3.8 Flash Next - 2 x R9700 vs. 3 x R9700 - 2 GPUs win
LocalLLaMA
MarcusAurelius68020
🟠 redditLiking Qwen Flash Next, what can I do for more speed?
LocalLLaMA
Forward_Jackfruit81329
🟠 redditQwen3.8-Flash-Next-oQ4e-mtp: 45 tok/s on M4 Max, 25 tok/s on M2 Ultra for local inference — llm-bench.io
LocalLLaMA
DerTomsn2917
🟠 reddit48 tg/s 440 prefill on my grandma's cluster (2xP40) (sort of)
LocalLLaMA
Jumpy-Operation-4615814
🟠 redditQwen 3.8 Next Flash is really really REALLY verbose..
LocalLLaMA
Infinite-Local54354373
🟠 redditMy Qwen3.8-27B task-aware quant reaches 99% of BF16 reasoning performance at 15% of the size.
LocalLLaMA
devildip21360
🟠 redditPrompting tips for 3.8 27B?
LocalLLaMA
Kahvana39
🟠 redditQwen3.8-Flash-Next on 2x3090: 9–12% faster decode at ~119k context, with a completed quality screen
LocalLLaMA
Extension-Bid-6392522
🟠 redditI made a custom llama.cpp build optimized for 7900xtx (one or two). for qwen 3.8 next and 27B. includes optimizations for PciE x4 and tensor parallel. read inside! (no AI slop)
LocalLLaMA
nasone325938
🟧 hnBenchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapsesstared268128
🟠 redditQwen 3.8 27b with PI agent - pushed to its 3D graphic game limits
LocalLLaMA
Healthy-Nebula-36036716
🟠 redditQwen 3.8 27b with PI agent - pushed to its 3D graphic game limits - locally
singularity
Healthy-Nebula-360342
🟠 redditQwen 3.8 27b with PI agent - pushed to its 3D graphic game limits
LocalLLaMA
Healthy-Nebula-360328798
🟠 redditQwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.
LocalLLaMA
FantasticNature75905523
🟠 redditQwen 3.8 27b with PI agent - pushed to its 3D graphic game limits - locally
singularity
Healthy-Nebula-36032811
🟠 redditQwen3.8-Flash-Next on MLX-serve, 1m context is released!
LocalLLaMA
Beamsters22255
🟠 redditWhat settings do you use for running Qwen3.8-Flash-Next in llama.cpp?
LocalLLaMA
HlddenDreck1348
🟠 reddit~1,400 t/s prefill is real. 60 t/s decode is not. I graded 5 Strix Halo forks with Qwen3.8 Flash-Next
LocalLLaMA
stereohype044
🟧 hnQwen 3.8 follows GPT-5.5 Pro reasoning prefillswsxiaoys23493
🟠 redditQwen3.8 Flash Next best quant fitting in 128GB Strix Halo the Mark Watney style
LocalLLaMA
MarkoMarjamaa422
🟠 redditQwen3.8-Flash-Next on 2x3090 + DDR4, part 4: 2.2-2.5x faster prefill by kicking the expert cache off the GPU while the prompt runs
LocalLLaMA
Extension-Bid-6392966
🟠 redditRunning Vision Qwen 3.8 27B on a 16GB Card, the config (45tks).
LocalLLaMA
FerLuisxd1829
🟠 redditQwen3.8 Flash Next UD-Q4_K_XL 49 tokens/s TGS using 2x RTX 3090 on Windows 11.
LocalLLaMA
whiteh4cker112
🟠 redditQwen3.8 Flash Next llama.cpp config tuning
LocalLLaMA
ChopSticksPlease7056
🟠 redditThis draft model is OP on 16 GB cards for Qwen 3.8 27b
LocalLLaMA
pneuny3714
🟠 reddit2×RTX 3090 + EPYC box running qwen3.8-flash-next at ~38 tok/s
LocalLLaMA
ludos1978937
🟠 redditQwen3.8 flash next - untrained svg generation
LocalLLaMA
ludos19788744
🟠 redditDear 24G owners, try VLLM you might be able to run Qwen3.8 27B INT4, 144K FP8 KV on RTX 3090 with better speed. (TLDR VLLM AOT)
LocalLLaMA
Altruistic_Heat_95318938
🟠 redditR9V Update: now ~100 tok/s in TG on Qwen3.8 Flash Next IQ4_XS on x2 R9700 + 128GB RAM. Fixed crashes with n-gram SSD streaming, improved diagnostics, plus pinned images. Q4_K_XL now supported, 50 tok/s TG.
LocalLLaMA
Public_Umpire_10995920
🟠 redditData point: Qwen3.8-Flash-Next PP/TG speed on M3 Ultra
LocalLLaMA
rm-rf-rm219
🟠 redditRunning Qwen3.8-Flash-Next locally on a 12GB VRAM card
LocalLLaMA
carteakey20175
🟠 redditI made a public Go-LLM guide for running Qwen3.8 Flash-Next with vLLM
LocalLLaMA
Motor_Ad1615
🟠 reddit2x5090 llama RPC Q2 Qwen Flash Next
LocalLLaMA
ilarp615
🟠 reddit5090 + r9700 in one case, configs inside: exl3, vllm nvfp4, and radiance mxfp4 running together
LocalLLaMA
IvGranite03
🟠 reddit[Release] SOTA GGUFs for Qwen3.8-Flash-Next: GSQ-RCO Providing Near Baseline Performance
LocalLLaMA
BullfrogScary89476946
🟠 redditQwen3.8-27B uncensored Q6_K at 156K context on one RTX 5090, 140-190 tok/s with DFlash2
LocalLLaMA
Fz1zz1810
🟠 redditQwen3.8 Flash on 12GB VRAM - 15 tokens/s
LocalLLaMA
KnownAd48324257
🟠 redditBenchmark: Qwen 3.8 27b NVFP4 on 2xV100 32gb, nvlink
LocalLLaMA
jjusko2027
🟠 reddit4xV620 Qwen3.8-flash-next 1300+ PP and 65+ TG on coding.
LocalLLaMA
Thin_Pollution884398
🟠 redditQwen3.8-Flash-Next on the official SGLang NVFP4 image: 35 s → 22 s first token at 254K on one RTX PRO 6000, then I run 7 tests
LocalLLaMA
FantasticNature759008
🟧 hnDwarf Star Support for Qwen3.8 Flash Nextmacote10
🟠 redditDoes MTP not work with ngram offloading in llamacpp for unsloth’s qwen3.8 next flash?
LocalLLaMA
Ambitious_Fold_287499
🟠 redditwhat's the state of the art recipe for running Qwen3.8-Flash-Next with a pair of 3090s and a ton of system RAM rn?
LocalLLaMA
starkruzr2032
🟠 redditBuilt this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark
LocalLLaMA
Character-Result-2815527
🟠 redditQwen3.8-Flash-Next on 5060 Ti looking for config advice
LocalLLaMA
MkGod529
🟠 redditis this good? 262k Qwen3.8:27B-Q4_K_M
LocalLLaMA
comperr214
🟠 redditQwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)
LocalLLaMA
peonist-ai12137
🟠 redditQwen-3.8-Flash-Next on 1x RTX 5090: TG=50 t/s, PP=2300 t/s - with FreeToken
LocalLLaMA
dir3ctly3729
🟠 redditOne more 'you should try ExllamaV3/exl3 for flash next' appreciation post
LocalLLaMA
youcloudsofdoom4542
🟠 redditQwen 3.8 next flash tuning
LocalLLaMA
Former-Tangerine-723032
🟧 hnQwen3.8-Flash-Next on a 64 GB M2 Ultra: A 66-Minute Real Work Runb1tank10
🟠 redditREAP/MTP Qwen3.8-Flash-Next MLX builds — here are the measured trade-offs to Jundot
LocalLLaMA
MensaProdigy1114
🟠 reddit5090 + 3090, Qwen3.8 27B and Flash-Next at 60k–160k context tests with llama.cpp
LocalLLaMA
Blindax1020
🟠 redditQwen3.8-27B: >70 tok/s (>160 tok/s concurrent), 10k tok/s prefill, full context on 2x3090 (or and 48GB or larger on ampere or higher), vanilla vllm
LocalLLaMA
maqifrnswa1932

Interpretation history

Decision trace