ISTA-DASLab β the Data Algorithms & Systems Lab at IST Austria β released GGUF quantizations of Alibaba's open-weights Qwen3.8-Flash-Next, a 512-expert MoE (~354GB at BF16), applying two of the lab's own methods: GSQ (described in coverage as Gumbel-Softmax Quantization) for per-tensor scalar quantization and RCO for budget-constrained per-tensor type allocation, with card-listed tiers at 2.40β3.50 bpw (66.4β83.6GB) claiming near-BF16 parity β exact AIME25 match at 3.00 bpw and the 3.50 IQ3_S tier 'matches or exceeds the base model on every task' (the card table shown here lists this tier, which the case had as title-only). A second, capability-targeted Coder build removes 50% of the 512 routed experts and quantizes the remainder to 3.5 bpw β 58.4GB total, an effective 1.89 bpw of the original β with the card claiming ~98.7% coding-performance retention. The snippets confirm both artifacts ship and their benchmark tables exist, but independent quality evidence within them is mixed: the lab's dense Qwen3.8-27B sibling quants were independently reported near-full quality at ~3 bpw, yet one tester found that sibling's smallest tier producing empty outputs, and the case history carries an unanswered negative coding hands-on of the extreme pruned Coder build β so 'usable at extreme compression' still rests mainly on the card's own benchmarks. Announced by u/Loginhe on r/Qwen_AI (~104 pts, ~30 days old), the release has since drawn field deployments via Strata expert streaming and a llama.cpp fork adding Q2_0 support specifically for the format.
ISTA-DASLab's per-tensor non-uniform GSQ-RCO quants independently arrive where Scott's own wiki already holds positions β per-tensor non-uniform quantization vs standard GGUF k-quants, and frontier-size MoEs runnable locally as sovereignty evidence β making this a dated receipt, not a repetition. Beyond convergence, the case bears directly on his open questions: the card's near-parity claims at 2.4β3.0 bpw versus duyntnet's unusable-for-coding ~1.89-bpw pruned-Coder hands-on is fresh evidence for exactly his 'minimum bpw before coding-agent quality degrades' and 'does half the expert pool matter' positions, the 66β76GB 3.0-bpw tier is directly testable on his 80GB-class gamepc box, and the Strata/ik_llama.cpp deployment ecosystem continues the expert-streaming and ByteShape/Bartowski per-tensor lineages the radar already tracks (radar:qwen38-flash-next-commodity-local-inference is a sibling episode of the same model, not a verdict on this case).
dev:concept.hardware-aware-local-inferencedev:project.gamepcwork:technology.large-language-modelsip:framework.sovereign-software-assuranceradar:qwen38-flash-next-commodity-local-inferenceradar:byteshape-qwen38-27b-quantsradar:bartowski-gguf-tensor-layoutsradar:concept.quantizationradar:concept.moe-inferenceradar:concept.expert-streamingradar:concept.extreme-quantization
queries asked of Scott's wikis
- per-tensor non-uniform quantization methods vs standard GGUF k-quants β Scott's position
- minimum bpw before coding-agent / agentic task quality degrades
- MoE expert redundancy and expert pruning β does half the expert pool matter
- expert offloading / CPU-GPU streaming for models oversized vs VRAM
- local-inference hardware envelope notes for 32β128GB consumer rigs
- open-weights sovereignty argument: frontier-size MoEs runnable locally
2026-10-10T00:28:47Z
Two new runability datapoints (IQ2_XS on RTX 3060 12GB; IQ3_S on 12GB VRAM via custom Strata fork) reinforce quantization practicality at extreme low bits, but capability uncertainty persists β no new quality evidence, no attribution resolution for hallucination/instruction-following failures at iq3βiq4. Case remains corroborated-but-cold pending systematic eval or sampling-param isolation.
2026-10-09T19:58:37Z
evidence attached: reddit.post.1x1mzsg β Additional user-reported performance for GSQ-RCO IQ3_S on 12GB VRAM via custom Strata fork; reinforces quantization practicality claims.
2026-10-09T19:58:37Z
evidence attached: reddit.post.1x1tclb β User-reported decode speeds for Qwen3.8-Flash-Next GSQ-RCO IQ2_XS on consumer hardware directly corroborates the open case's quantization claims.
2026-10-08T07:29:19Z
New Q2_0 vs IQ2_XS Strata comparison adds another extreme-low-bit runability datapoint, but the velocity spike (86obsessed thread 135/239) is repetitive amplification of the already-priced iq3βiq4 quality contest β no new capability evidence, no attribution resolution. Scott's down-vote stands; world hasn't moved. Case stays corroborated-but-cold pending systematic eval or sampling-param isolation.
2026-10-08T07:01:41Z
evidence attached: reddit.post.1x0j0ae β User compares Q2_0 vs IQ2_XS GSQ-RCO quantizations of Qwen3.8-Flash-Next using Strata runtime, directly engaging with the quantization variants claimed in the case.
2026-10-07T17:18:30Z
No new facts this look: the activity is growth inside the already-priced iq3βiq4 quality debate (86obsessed thread now 95/186, WishfulAgenda adds a same-tier negative, same confounds), i.e. repetitive amplification of a recorded contest, not peripheral expansion β the 20x velocity spike and 91st-percentile peer rate are volume without substance, so heat stays low despite the speedometer. Scott's explicit not-relevant down-vote is respected (the world has not genuinely moved): relevance drops highβlow and the case settles into corroborated-but-cold pending an attribution-settling eval or a dev/user response isolating sampling params.
2026-10-07T05:00:27Z
The contested zone widened: multi-user negative hands-ons now cover the mid iq3βiq4 MoE tiers in agentic work (32% hallucinated site index at iq3; instruction-following failures at iq4_xs), not just the extreme IQ1 Coder build β though live confounds (Strata server sampling params, Flash-Next's own preview-model glitchiness) blur how much loss is quant-attributable. The case now reads: runability and the prune/quant route proven, dense-sibling quality confirmed at method level, while the MoE's own capability retention is actively contested in the community's most-discussed recent thread.
2026-10-07T03:33:13Z
evidence attached: reddit.post.1wzj990 β Community counter-evidence that Flash-Next degrades vs 27B dense at iq3βiq4 quants in agentic work (32% hallucinated site index at iq3), directly qualifying the extreme-compression usability claim.
2026-10-07T01:36:09Z
grounded: converges/high β ISTA-DASLab's per-tensor non-uniform GSQ-RCO quants independently arrive where Scott's own wiki already holds positions β per-tensor non-uniform quantization vs
2026-10-07T01:26:21Z
The promotion blocker β independent capability confirmation β is now satisfied at method level: two same-day unrelated users report GSQ-RCO IQ3_S on Qwen3.8-27B delivering near-full-precision quality at ~11β16GB, joining the proven deployment ecosystem (multiple rigs, Strata consensus, ik_llama.cpp fork) as a second independent evidence line. The case's contested residue has narrowed sharply to the extreme tiers (2.4 bpw MoE builds, IQ1/Q2 pruned Coder) where duyntnet's negative coding hands-on stands unanswered; the route is corroborated, the extreme-compression clause is not.
2026-10-06T23:36:36Z
evidence attached: reddit.post.1wzfv45 β Second same-day independent report of GSQ-RCO IQ3_S near-full-precision quality (16GB build with KV streaming to 262k), corroborating the quant family's retention claim.
2026-10-06T23:36:36Z
evidence attached: reddit.post.1wzha67 β Independent user reports GSQ-RCO IQ3_S runs Qwen3.8-27B near full-precision quality at ~11GB on stock llama.cpp β method-level corroboration of the quant family's capability-retention claim.
2026-10-06T04:36:31Z
IceFog72's ik_llama.cpp fork is the case's first toolchain adoption β quant-type support (Q2_0) written specifically for GSQ-RCO artifacts, with a working IQ1_M Coder run β meaning the format has crossed from things-people-run to infrastructure-people-build-on. But this deepens the already-proven runability half without touching the contested capability half (duyntnet's negative hands-on still stands unanswered, no independent evals), so the case holds at watching/low; promotion now waits specifically on independent quality confirmation.
2026-10-06T03:33:52Z
evidence attached: reddit.post.1wynomx β Fork adds Q2_0 support specifically to run GSQ-RCO quants, with a working IQ1_M run of Flash-Next-Coder in comments β adoption evidence for the extreme-compression route.
2026-10-04T21:45:34Z
A third independent deployment β a user running the directly-linked ISTA-DASLab GSQ-RCO-Coder GGUF (~27.6GB, IQ1-class) on 2x 5060 Ti + 32GB RAM β upgrades the previously title-only pruned Coder build from 'unverified' to 'ships and runs, quality unassessed', and the same thread's consensus that Strata (not llama.cpp) is the way to run these quants consolidates the deployment pattern. Capability remains contested with no evals, and at 0.17 pts/h (39th percentile, steady) the numbers agree with low heat, but the compression-envelope story now demonstrably outperforms the claimed 58β84GB floor.
2026-10-04T21:24:13Z
evidence attached: reddit.post.1wxpvys β Third-party user actually deploying a GSQ-RCO IQ1 quant of Flash-Next at ~27.6GB on consumer hardware β real-world adoption evidence beyond the claimed 58-84GB envelope.
2026-10-02T20:17:09Z
The case's meaning shifts from 'released and benchmarked' to 'released, benchmarked, and field-run': Fz1zz's deployment of the case's own GSQ-RCO IQ3_XXS at full 262K context on 31GB RAM + 48GB VRAM via Strata expert streaming (130 tok/s decode) supplies the first independent decode numbers and verifies Strata by use β but the deployment thread's commenters ask the quality question and no one answers, so usable-capability stays contested against duyntnet's negative hands-on, and the 1.89 bpw pruned Coder build remains title-only. The practical-envelope half of the hypothesis is now independently demonstrated; the capability half is not.
2026-10-02T19:29:57Z
evidence attached: reddit.post.1ww2lv9 β Field deployment runs the case's own GSQ-RCO IQ3_XXS quant at full 262K context on just 31GB RAM via Strata expert streaming, pushing the practical envelope below the claimed 58β84GB.
2026-10-02T10:59:32Z
The velocity-spike triggers are multiples of near-zero baselines (3β8 pts/h absolute on the low-traffic Victoria post, against 0.08β1.0 peer baselines) β sensor noise, not renewed attention; measured velocity is 0 pts/h at the 11th peer percentile on a ~620h-old case. No new evals, no verification of the title-only 1.89 bpw pruned Coder build, no Strata repo content arrived, so the case keeps its shape: GSQ-RCO usability still contested (card parity claims vs duyntnet's unusable-for-coding hands-on), prune-and-retrain route independently proven by Victoria, patient watching brief unchanged.
2026-10-01T00:56:43Z
The prune-and-quantize half gains independent route-level proof-of-concept on this exact MoE: rmonsurate's Victoria (REAP 44% expert cut + NVFP4 retrain, GGUF shipped, 70% TB2.1) shows expert removal with retraining retains coding/agent capability β but by a different method, so it strengthens the route without validating the still title-only 1.89 bpw 50%-pruned Coder build. With the GSQ-RCO quants' real-world usability still contested (near-parity benchmarks vs duyntnet's unusable-for-coding hands-on) and traffic cooled to ~1 pt/h from a ~41 pt/h peak, this becomes a patient watching brief pending systematic evals rather than an urgent look.
2026-10-01T00:29:45Z
evidence attached: reddit.post.1wujph3 β Fuller post of the same Victoria/Maple release with GGUF results (75.3% TB) further corroborating the prune-and-quantize route on this model.
2026-10-01T00:29:45Z
evidence attached: reddit.post.1wujtbj β REAP expert-pruning plus NVFP4 retrain of Qwen3.8-Flash-Next holding 70% TB2.1 is independent corroboration that expert-pruning retains agentic capability on this MoE.
2026-09-30T13:39:24Z
The case's meaning shifted from promising first-party claims to a contested live test: third-party tooling (Strata) already claims support with 'high tps on consumer hardware' β the first independent adoption line β while the first hands-on report (duyntnet) found the quants unusable for coding (hallucinated Delphi/C++ syntax vs working 27B quants), directly contesting the 'usable capability' half and putting the ~1.89 bpw 50%-pruned Coder build most at risk. The velocity spike has cooled to ~0/h, so this is now a watching brief pending systematic community evals rather than an urgent look.
2026-09-29T09:46:39Z
origin walked (opencode/cheap-glm, conf 0.86): anchor reddit.post.1wt4s88 -> echo.other.5778328b4d by ISTA-DASLab (Deep Algorithms and Systems Lab, Institute of Science and Technology Austria)
2026-09-29T09:33:52Z
grounded: novel/high β First-party ISTA-DASLab release lands squarely on Scott's hardware-aware local-inference territory: a 512-expert 176.9B MoE at 66.4β75.8GB with 99.4% BF16 recov
2026-09-29T09:25:17Z
case created β A first-party release proposing a distinct joint quantize-and-prune method for a very large MoE, separate from the ByteShape and Bartowski episodes, and resolvable by community benchmarking and adoption.