2026-10-11 17:21 UTC

Qwen claims its open-weight Qwen3.8-Flash-Next model uses a reworked hybrid-attention architecture to make long-context and agentic inference more efficient, potentially expanding practical local deployment.

state: resolvedheat: lowuncertainty: mediumconvergesscott: highqwen38-flash-next open-models local-inferenceQwen

What is this?

Qwen, identified in the snippets as Alibaba’s model team, has staged Qwen3.8-Flash-Next on Hugging Face as an upcoming open-weight preview of the Qwen4 architecture. The supplied results describe a hybrid-attention, multimodal mixture-of-experts design intended to improve long-context efficiency, but they do not provide an official architecture specification or benchmarks. Claims that it will be practical for local or agentic inference remain prospective and partly community-sourced; one related architecture report still estimates substantial hardware requirements.

Why it matters to Scott

Qwen’s open-weight, efficiency-oriented design converges with Scott’s active hardware-aware local-inference work and could become a practical option for his self-hosted GPU substrate. The impact remains conditional because the supplied evidence lacks official specifications, benchmarks, and proof that the model runs usefully on consumer hardware.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:framework.sovereign-software-assuranceradar:concept.open-modelsradar:concept.local-inferenceradar:concept.moe-inferenceradar:concept.long-context-inferenceradar:longcat-sparse-24gb-inferenceradar:dwarf-sparse-attention-validation
queries asked of Scott's wikis
  • hybrid attention for long-context agents
  • sparse MoE local inference economics
  • open-weight models and model sovereignty
  • local multimodal agent deployment
  • long-context efficiency versus agent memory
  • architecture-aware inference harnesses

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (24) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit[Megathread] Qwen3.8-Flash-Next - Release Day
LocalLLaMA
sammcj430593
🟧 echo.blog ⭐The official Qwen announcement says: “In this release we are opening the weights of Qwen3.8-Flash-Next,” describing it as a multimodal MoE aQwen Team——
🟠 redditDeepSeek V4 0731 -> Qwen 3.8 Flash -> GLM 5.3 Flash (and back again!)
LocalLLaMA
Legitimate_Hat_78522740
🟠 redditQwen3.8-Flash-Next: Time to Update Those Benchmarks
LocalLLaMA
tolitius15437
🟠 redditHas anybody got Qwen3.8 Flash to work on 2x DGX Sparks?
LocalLLaMA
StartupTim1013
🟠 redditfeat: import qwen4exp (Qwen3.8-Flash-Next) support from upstream PR #27742 by giveen · Pull Request #324 · TheTom/llama-cpp-turboquant
LocalLLaMA
giveen110
🟠 redditllama.cpp support for Qwen3.8-Flash-Next has been merged
LocalLLaMA
jacek202337292
🟠 redditIs Qwen3.8-Flash-Next really fast (or am I doing something wrong with Qwen3.6)?
LocalLLaMA
jtabernik00
🟠 redditQwen 3.8 Flash Next: Ali Baba ditches Apache 2.0 licence
LocalLLaMA
Mr_Moonsilver3719
🟠 redditQwen3.8-Flash-Next much worse than DeepSeek-V4-Flash?
LocalLLaMA
vini542reddit030
🟠 redditInitial thoughts on 3.8 Next IQ3
LocalLLaMA
Repulsive_Initial308017
🟠 redditBest parameter setting for Qwen3.8 Flash Next on llama CPP
LocalLLaMA
Motor_Ad16421
🟠 redditWhat is Qwen 3.8 Next Engram usage?
LocalLLaMA
I-am_Sleepy1711
🟠 redditQwen3.8-Flash-Next (UD-IQ4_XS) on 2x RTX 3060 + 7800X3D, from initial 36 tps prefill to 400 tps and other benchmarks (-sm tensor trap) + VRAM/RAM usage
LocalLLaMA
Comrade_Mugabe6628
🟠 redditIt's unbelievable! I used the mmap function in llama.cpp to fit Qwen3.8-Flash-Next IQ3_XSS into 16G+64G RAM, and the speed still reached 26t/s.
LocalLLaMA
q801922211267
🟠 redditTQwen 3.8 flash next ud1s on 6gb vram and 16 gb system ram
LocalLLaMA
Agitated_Force_91991319
🟠 redditQwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM)
LocalLLaMA
crusaderky14838
🟠 redditQwen3.8-Flash-Next FP8 running at 524K context on 2x RTX PRO 6000 with vLLM — found an MTP long-context issue
LocalLLaMA
SpendLucky12731020
🟠 redditToday I hit 181 toks/s (aggregate) on Qwen3.8-Flash-Next on 2x DGX Sparks
LocalLLaMA
StartupTim23056
🟠 redditIs it worth running Qwen 3.8 Flash Next on 4x3090 vs 27B?
LocalLLaMA
Acceptable_Adagio_915277
🟠 redditQwen 3.8 Flash Next ngram look up table offloaded to SSD and streamed in SGLang
LocalLLaMA
Easy_Werewolf790335157
🟠 redditAtomicChat/Qwen3.8-Flash-Next-GGUF is Really Good
LocalLLaMA
tolitius6313
🟠 redditQwen3.8-Flash-Next + MTP on Strix Halo: Vulkan Runtime Notes
LocalLLaMA
betiz02610
🟠 redditQwen3.8-Next streaming - 150tps prefill, 3.6 tps decode on M5 Air
LocalLLaMA
maddie-lovelace189

Interpretation history

Decision trace