2026-10-11 16:38 UTC

Loginhe claims the released GSQ-RCO quantizations (2.40–3.50 bpw) plus a 50%-expert-pruned Coder build of the 176.9B-parameter Qwen3.8-Flash-Next MoE retain usable capability at extreme compression, establishing quantization-plus-expert-pruning as a practical route to running very large MoE models in roughly 58–84GB of memory.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: lowquantization expert-pruning local-inferenceLoginhe

What is this?

ISTA-DASLab β€” the Data Algorithms & Systems Lab at IST Austria β€” released GGUF quantizations of Alibaba's open-weights Qwen3.8-Flash-Next, a 512-expert MoE (~354GB at BF16), applying two of the lab's own methods: GSQ (described in coverage as Gumbel-Softmax Quantization) for per-tensor scalar quantization and RCO for budget-constrained per-tensor type allocation, with card-listed tiers at 2.40–3.50 bpw (66.4–83.6GB) claiming near-BF16 parity β€” exact AIME25 match at 3.00 bpw and the 3.50 IQ3_S tier 'matches or exceeds the base model on every task' (the card table shown here lists this tier, which the case had as title-only). A second, capability-targeted Coder build removes 50% of the 512 routed experts and quantizes the remainder to 3.5 bpw β€” 58.4GB total, an effective 1.89 bpw of the original β€” with the card claiming ~98.7% coding-performance retention. The snippets confirm both artifacts ship and their benchmark tables exist, but independent quality evidence within them is mixed: the lab's dense Qwen3.8-27B sibling quants were independently reported near-full quality at ~3 bpw, yet one tester found that sibling's smallest tier producing empty outputs, and the case history carries an unanswered negative coding hands-on of the extreme pruned Coder build β€” so 'usable at extreme compression' still rests mainly on the card's own benchmarks. Announced by u/Loginhe on r/Qwen_AI (~104 pts, ~30 days old), the release has since drawn field deployments via Strata expert streaming and a llama.cpp fork adding Q2_0 support specifically for the format.

Why it matters to Scott

ISTA-DASLab's per-tensor non-uniform GSQ-RCO quants independently arrive where Scott's own wiki already holds positions β€” per-tensor non-uniform quantization vs standard GGUF k-quants, and frontier-size MoEs runnable locally as sovereignty evidence β€” making this a dated receipt, not a repetition. Beyond convergence, the case bears directly on his open questions: the card's near-parity claims at 2.4–3.0 bpw versus duyntnet's unusable-for-coding ~1.89-bpw pruned-Coder hands-on is fresh evidence for exactly his 'minimum bpw before coding-agent quality degrades' and 'does half the expert pool matter' positions, the 66–76GB 3.0-bpw tier is directly testable on his 80GB-class gamepc box, and the Strata/ik_llama.cpp deployment ecosystem continues the expert-streaming and ByteShape/Bartowski per-tensor lineages the radar already tracks (radar:qwen38-flash-next-commodity-local-inference is a sibling episode of the same model, not a verdict on this case).
dev:concept.hardware-aware-local-inferencedev:project.gamepcwork:technology.large-language-modelsip:framework.sovereign-software-assuranceradar:qwen38-flash-next-commodity-local-inferenceradar:byteshape-qwen38-27b-quantsradar:bartowski-gguf-tensor-layoutsradar:concept.quantizationradar:concept.moe-inferenceradar:concept.expert-streamingradar:concept.extreme-quantization
queries asked of Scott's wikis
  • per-tensor non-uniform quantization methods vs standard GGUF k-quants β€” Scott's position
  • minimum bpw before coding-agent / agentic task quality degrades
  • MoE expert redundancy and expert pruning β€” does half the expert pool matter
  • expert offloading / CPU-GPU streaming for models oversized vs VRAM
  • local-inference hardware envelope notes for 32–128GB consumer rigs
  • open-weights sovereignty argument: frontier-size MoEs runnable locally

Measured heat

now 0 pts/hpeak 119 pts/hcomments 0/hpeers p25momentum: steady2 platformsage 842h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-06 14:00⭐ origin echo-reconstructedThe HF model card is the primary artifact the Reddit post announces: "This repository provides GGUF quantizations of Qwen3.8-Flash-Next at t
ISTA-DASLab (Deep Algorithms and Systems Lab, Institute of Science and Technology Austria) on other (echo) Β· attributed from reddit.post.1wt4s88
β€”
09-29 08:40first on r/LocalLLaMA Β· published Β· +546.7h[Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw
Loginhe
β€”
09-29 08:40amplified on r/LocalLLaMAreddit.post.1wt4s88
Loginhe
peak 104 Β· 77 comments Β· 16% of case engagement
09-30 23:10amplified on r/LocalLLaMAreddit.post.1wujph3
rmonsurate
peak 110 Β· 40 comments Β· 14% of case engagement
09-30 23:14amplified on r/LocalLLaMAreddit.post.1wujtbj
rmonsurate
peak 1 Β· 9 comments Β· 1% of case engagement
10-02 19:19amplified on r/LocalLLaMAreddit.post.1ww2lv9
Fz1zz
peak 8 Β· 15 comments Β· 2% of case engagement
10-04 20:35amplified on r/LocalLLaMAreddit.post.1wxpvys
Bulky-Priority6824
peak 1 Β· 15 comments Β· 1% of case engagement
10-05 23:33amplified on r/LocalLLaMAreddit.post.1wynomx
IceFog72
peak 35 Β· 7 comments Β· 4% of case engagement
6 more amplifiers in ainews.case_chain
09-29 09:20our radar first saw it Β· +547.3hdiscovery anchor: reddit.post.1wt4s88β€”
pace: p81 vs 519 stories at the 720h mark (now 842h old) β€” ahead of nvidia-sol-pi-harness-efficiency (1.1x), behind coop-coding-agent-vm-isolation (1.0x)

Evidence (13) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit[Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw
LocalLLaMA
Retrieved article excerpt

Open article Β· Retrieved 2026-09-29T09:24:00.036696+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. Β© "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
Loginhe10477
🟧 echo.other ⭐The HF model card is the primary artifact the Reddit post announces: "This repository provides GGUF quantizations of Qwen3.8-Flash-Next at tISTA-DASLab (Deep Algorithms and Systems Lab, Institute of Science and Technology Austria)β€”β€”
🟠 redditReleasing two post quantized trained models with fine tuned draft heads: Victoria (REAP + QAD) 44% smaller Qwen Flash Next, 70% TB 2.0 avg of three runs & Maple: Victoria + Canandian dataset finetune
LocalLLaMA
rmonsurate09
🟠 redditTwo open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune)
LocalLLaMA
rmonsurate11040
🟠 redditQFN at 262K on 32 GB RAM 48GB VRAM : 1,800 tok/s prompt, 130 tok/s decod three Strata patches.
LocalLLaMA
Fz1zz615
🟠 redditQFN llama.cpp Any juice left to squeeze?
LocalLLaMA
Bulky-Priority6824011
🟠 redditk_llama.cpp MoE Optimizations: Expert Residency, Hybrid CPU/GPU Execution, Q2_0 Support
LocalLLaMA
IceFog72307
🟠 redditQwen3.8-27B with just 12GB VRAM + 8GB RAM - ~18 tok/sec
LocalLLaMA
bodhi3711712
🟠 redditjust my "how I run qwen3.8 27b on 16GB" experience and guide
LocalLLaMA
randomgenericbot224
🟠 redditUgh I didn't want to post this... Back to Qwen3.8 27B
LocalLLaMA
86obsessed131243
🟠 redditQwen Flash Q2_0 vs IQ2_XS GSQ-RCO using Strata
LocalLLaMA
dampflokfreund06
🟠 redditQwen 3.8 Flash Next-GSQ-RCO-IQ2_XS at ~21 tok/s on just an RTX 3060 12GB + 16GB DDR4 RAM(No gate pruning, 100% bit-exact)
LocalLLaMA
zyxciss13034
🟠 redditQwen3.8-Flash-Next-GSQ-RCO (IQ3_S): ~20-30 tok/sec decode & 300-90k tok/sec prefill on 12GB VRAM + 32GB RAM + NVME
LocalLLaMA
bodhi3714730

Interpretation history

Decision trace