2026-10-11 18:04 UTC

Community operators and quant maintainers claim optimized RAM/NVMe offload and compact GGUF quants make Qwen3.8-Flash-Next practically runnable on commodity systems ranging from one 12GB GPU to dual RTX 3090s.

state: resolvedheat: lowuncertainty: lowknownscott: mediumqwen local-inference model-offloading quantizationQwenllama.cppvLLMagentionai
Surfaced 2026-09-01T06:28:16Z โ€” priced heat=high at reprice: MTP has crossed from a downloadable but unusable GGUF artifact into reportedly merged llama.cpp support, with initial operator measurements showing substantial decode gains. This closes the prior runtime-support gap, though commodity-system reproduction, quality effects, and long-context performance still need validation.

What is this?

Qwen3.8-Flash-Next is described as a Qwen open model with a 125B-parameter mixture-of-experts architecture activating roughly 6B parameters per token, alongside a large n-gram component. Community posts and quant titles claim that low-bit GGUF builds plus expert and n-gram offloading to system RAM or NVMe make it runnable on hardware ranging from one 12GB GPU to dual RTX 3090s, while vLLM support is reported and llama.cpp optimization is underway. The evidence is preliminary and conflicting: an NVIDIA forum thread calls the specifications a rumor, Unsloth says its support is still in progress, and the supplied snippets do not independently verify the claimed single-GPU performance.

Why it matters to Scott

The radar already tracks this mechanism and validation question in `radar:deepseek-v4-nvme-demand-paging`, `radar:hotpin-lossless-moe-streaming`, `radar:layerstorm-moe-expert-streaming`, and `radar:longcat-sparse-24gb-inference`; this is another preliminary model-specific claim rather than a new position. It still bears directly on Scottโ€™s `gamepc` model-serving substrate and hardware-aware local inference policy, where verified RAM/NVMe offload performance could expand the models practical on his existing hardware.
dev:project.gamepcdev:concept.hardware-aware-local-inferenceradar:deepseek-v4-nvme-demand-pagingradar:hotpin-lossless-moe-streamingradar:layerstorm-moe-expert-streamingradar:longcat-sparse-24gb-inference
queries asked of Scott's wikis
  • heterogeneous RAM NVMe offload for local inference
  • low-bit quantization quality versus hardware accessibility
  • MoE active parameters versus model storage economics
  • commodity hardware strategy for frontier-scale open models
  • local model sovereignty and private agent workloads
  • llama.cpp and vLLM inference harnesses

Measured heat

no measured readings yet โ€” the hourly heat pass fills this in

How the heat travelled

no chain yet โ€” the hourly chain pass fills this in

Evidence (58) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญRunning Qwen3.8-Flash-Next (125B MoE + 51B n-gram table) on 2ร—RTX 3090 + 96GB DDR5. Optimised Llama.cpp and vLLM with experts offload to RAM and n-gram offload to NVME (proven and being optimised)
LocalLLaMA
jbro1985243
๐ŸŸ  redditQwen3.8 Flash Quants
LocalLLaMA
Dutchnamn1815
๐ŸŸ  redditQwen3.8-Flash-Next IQ1_S on a single 5070 (12GB VRAM)
LocalLLaMA
jacek2023617
๐ŸŸ  redditIโ€™ve pushed llama.cpp pretty far for Qwen3.8-Flash-Next โ€” is there any reason not to move to vLLM for 200K+ context?
LocalLLaMA
Prudent_Appearance711020
๐ŸŸ  redditqwen3.8-flash-next and 262144 vs 1M context - RoPE/YaRN and llama-server built with PR 27742
LocalLLaMA
burritoresearch10
๐ŸŸ  redditQwen3.8-Flash-Next protip tensor-read-lazy on requires load-mode mmap
LocalLLaMA
Pyrolistical63
๐ŸŸ  redditAny reason to use Qwen 3.8 Next UD IQ1_S over Qwen 3.8 27B UD Q4?
LocalLLaMA
freedomachiever525
๐ŸŸ  redditRan Qwen3.8-Flash-Next (79 GB, 2-bit) at 350K ctx for 3.5 hours on a 128 GB M5 Max โ€” speed vs context depth, 100 turns, one graph
LocalLLaMA
Artistic_Okra72884412
๐ŸŸ  redditQwen3.8-Flash-Next optimised for Macs
LocalLLaMA
memeka4825
๐ŸŸ  redditQwen3.8-Flash-Next NVFP4 Day-3 support for 4xV100
LocalLLaMA
Primary_Exchange21337
๐ŸŸ  redditQwen3.8-Flash-Next at 170K context on a single 96 GB card. ~110 tok/s.
LocalLLaMA
UltrMgns2123
๐ŸŸ  redditQwen 3.8 Flash Next Q6_K_XL loads OK on 4 x 3090.
LocalLLaMA
ethertype59
๐ŸŸ  redditExperience report - Qwen 3.8 Flash Next on memory rich, GPU poor setup
LocalLLaMA
Positive-Stock64445445
๐ŸŸ  redditQwen3.8-Flash-Next NVFP4 2xDGX Spark config: 50t/s decode, 2,900t/s prefill
LocalLLaMA
-dysangel-2327
๐ŸŸ  redditDon't Sleep on EXL3 Quants
LocalLLaMA
PyaesoneP3949
๐ŸŸ  redditGuys, if You are Starved for RAM to Run Qwen3.8-Next-Flash at Q4, Try the Atomic Chat: It's Highly Memory Efficient and Fast!
LocalLLaMA
Iory1998044
๐ŸŸ  redditPi Agent - Using Qwen3.8 Flash Next Q4_K_XL on the Little Man RTX, Qwen3.8 27B Q8_K_XL on the Chonky Boi W7900 - Pi treats the 27B as an Oracle in co-development
LocalLLaMA
Thrumpwart43
๐ŸŸ  redditBest settings for harness work with llama.cpp + qwen 3.8
LocalLLaMA
GodComplecs57
๐ŸŸ  redditQwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM
LocalLLaMA
sloptimizer3327
๐ŸŸ  redditQwen 3.8 Flash Next locally on simple mobile phone at 3.5 tok/s
LocalLLaMA
dai_app7045
๐ŸŸ  redditQwen Flash Q4_K_M on 4080 + 64GB DDR5 at ~8tk/s 98304 CTX.
LocalLLaMA
Desperate-Data-37471018
๐ŸŸ  redditR9V: A designer set of kernels I've been working on for R9700s/RDNA4. Qwen3.8-Flash-Next Unsloth IQ4_XS (w/ TP on 2 R9700s, MTP, SSD n-gram, 128k ctx, vision): TG256 of *78 tok/s* (~3x increase), PP8192 of *1510 tok/s* (~30x increase).
LocalLLaMA
Public_Umpire_10992222
๐ŸŸ  redditQwen3.8-Flash-Next-NVFP4 vs Qwen3.8-27B-FP Test Results
LocalLLaMA
trashacct3836238
๐ŸŸ  redditQwen3.8-Flash-Next BF16 at 6.4 t/s with RTX PRO 6000 64GB DDR5, RTX 5090 64GB DDR5, MacBook Pro 48GB
LocalLLaMA
stargate425011
๐ŸŸ  redditQwen3.8-Flash-Next on a 96GB Mac Studio (here's my memory math, tell me where it's wrong)
LocalLLaMA
Mxmtm510
๐ŸŸ  redditHow I got Qwen 3.8 27b running at ~75t/s decode on 16GB RTX 5080
LocalLLaMA
Kernoriordan6981
๐ŸŸ  redditQwen3.8-Next-Flash up to 240t/s on single rtx 6000 pro
LocalLLaMA
AdventurousSwim1312724
๐ŸŸ  redditQwen3.8 27B Q3S just created this and thought of all the necessary features. im just blown away. so cool.
LocalLLaMA
Old-Sherbert-449504
๐ŸŸ  redditMy Qwen3.8-Flash-Next recipe for single GB10/DGX Spark, uses Intel AutoRound int4 quant and vLLM, fp8 ngram table offloaded to local SSD or external RDMA server. At mtp=3 c=1, code is ~47.5t/s, json is ~60t/s. Prefix cache is ON.
LocalLLaMA
Saren-WTAKO61
๐ŸŸ  redditThe Chrono Trigger plot challenge - Crono awakens in his modest bedroom of 2095...
LocalLLaMA
rpwoerk2619
๐ŸŸง hnShow HN: Slotstream, run Qwen3.8-Flash-Next 4-bit on a low-memory Maccarloslfu20
๐ŸŸ  redditCan I get more out of Qwen 3.8 Flash Next (4-bit - 64gb VRAM + 32gb RAM)
LocalLLaMA
Jorlen138
๐ŸŸง hnThe Llama.cpp Fork That Enables Qwen 3.8 27B Large Contexts for 16GB VRAM GPUdazhbog42
๐ŸŸ  redditQwen 3.8 on 3090 comparisons
LocalLLaMA
LittleCelebration412032
๐ŸŸ  redditWhats the current state of Qwen 3.8 Flash regarding inference (llama.cpp)?
LocalLLaMA
No_Algae17533771
๐ŸŸ  redditQwen3.8-Flash-Next in llama.cpp from CPU-only to 96GB VRAM: 8.5 to 109 tok/s, max context and parameters test. My findings on RTX 6000 PRO.
LocalLLaMA
FantasticNature75906029
๐ŸŸ  redditWarning: llama.cpp --lazy-mode default changed to auto - large tables may stay on disk
LocalLLaMA
whiteh4cker3823
๐ŸŸ  redditSadly, there are no good Qwen3.8 27B NVFP4 GGUF
LocalLLaMA
Pyrolistical029
๐ŸŸ  redditQwen 3.8 Flash Next Benchmarks on TensorSharp and llama.cpp
LocalLLaMA
fuzhongkai00
๐ŸŸ  redditMTP released for Qwen3.8-Flash-Next-GGUF
LocalLLaMA
vini542reddit457103
๐ŸŸ  redditExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++
LocalLLaMA
Unstable_Llama175134
๐ŸŸ  redditqwen4exp fixes in llama.cpp
LocalLLaMA
jacek20235118
๐ŸŸ  redditWhich current local models that can run within 128GB generate the best SVG pelicans?
LocalLLaMA
pmigdal4164
๐ŸŸ  redditI pushed Qwen3.8-27B to 2.000 prefill per second and 132 decode per second on A RTX 3090.
LocalLLaMA
iamMess11779
๐ŸŸ  redditUpdate: llama.cpp for Radeon VII / MI50 / MI60 โ€” +14% PP, +9% long-context fill vs upstream + adaptive Flash Attention
LocalLLaMA
milpster226
๐ŸŸง hnShow HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/scarloslfu21397
๐ŸŸ  redditAnyone else notice strange refusal-related reasoning traces from Qwen3.8-Flash-Next during routine coding sessions?
LocalLLaMA
wombweed65
๐ŸŸ  redditHow I got 280 tok/s on Qwen3.8 27B on 2xr9700's and 940k tokens kv cache
LocalLLaMA
whodoneit16823
๐ŸŸ  redditQwen3.8-Flash-Next (104 GB MoE) on a Strix Halo + RTX 3090 Ti eGPU: 22 -> 84 tok/s, and within one HumanEval+ problem of a dual-3090 vLLM box at 0.4x the wall time
LocalLLaMA
TrifleHopeful541879
๐ŸŸ  redditRunning 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s
LocalLLaMA
yogthos9816
๐ŸŸ  redditi unlocked P2P on two 5060ti but failed
LocalLLaMA
chocofoxy411
๐ŸŸ  redditConfirmed bolting Q8 NGram into IQ4 Qwen no speed degradation
LocalLLaMA
Altruistic_Heat_95317820
๐ŸŸ  redditMTP or MTP+Ngram for Qwen3.8 Flash Next?
LocalLLaMA
esw12306
๐ŸŸ  redditQwen3.8 Flash AP Quants
LocalLLaMA
Dutchnamn1110
๐ŸŸ  redditQwen3.8-flash-next sees corruption everywhere
LocalLLaMA
arkham002234
๐ŸŸ  redditQwen 3.8 users (flash next and 27b) - do you force reasoning to low? Better results that way?
LocalLLaMA
Jorlen044
๐ŸŸ  redditQwen3.8-Flash-Next on 2x3090 + DDR4: 17 โ†’ 25-29 t/s decode with the expert cache PR
LocalLLaMA
Extension-Bid-6394955
๐ŸŸง hnI ran Qwen3.8-27B locally on an RTX 5060 Ti 16GBarczhi20

Interpretation history

Decision trace