2026-10-11 16:38 UTC

ByteShape claims its released ShapeLearn Qwen 3.8 27B GGUFs retain 99.63% of BF16's aggregate eight-benchmark score at 3.84 bits per weight and improve its measured quality-speed-memory frontier, potentially improving practical local-model deployment tradeoffs.

state: watchingheat: lowuncertainty: mediumconvergesscott: mediumquantization local-inference inference-economics qwenByteShapeenrique-byteshape
Surfaced 2026-09-19T08:26:03Z β€” priced heat=high at reprice: Independent AMD deployment reports now qualify the speed story: users report no advantage over Unsloth on Vulkan and substantially slower draft-assisted decoding than MTP, alongside usable MTP performance. Expanding Reddit/Hacker News attention and concrete hardware trials warrant high heat, but neither the benchmark-retention claim nor a general quality-speed-memory frontier improvement is independently corroborated.

What is this?

ByteShape, a model-optimization shop that also publishes quantized image-generation builds, released five per-tensor 'ShapeLearn' GGUF quantizations of Alibaba's open Qwen 3.8 27B, replacing its earlier ShapeLearn-Lite quants. Its headline claims are self-measured: GPU-5 retains 99.63% of the BF16 model's aggregate score across an eight-benchmark suite (math, coding, knowledge, instruction following, agentic tool use) at 3.84 bits per weight (13.1 GB), GPU-4 98.72% at 3.23 bpw (11.0 GB), with speculative-decoding speedups (MTP to ~1.57x, DFlash2 to ~1.78x on an RTX 5090) also self-reported. Community evidence in the supplied snippets qualifies rather than confirms this: an HN user reports no speed advantage over Unsloth on a bandwidth-bound 7900 XTX (~60 t/s), and commenters dispute circulation framing, while the case file records independent low-bit quality and draft-decoding complaints on variants other than GPU-5. Nothing in the snippets constitutes independent replication of the 99.63% aggregate; all quality figures trace back to ByteShape's own BF16-normalized evaluation.

Why it matters to Scott

The episode's independent-test resolution β€” vendor aggregate-retention claims that don't replicate, speed wins that vanish off NVIDIA, and per-workload inconsistency behind a strong BF16-normalized aggregate β€” converges with dev:concept.hardware-aware-local-inference's core position that precision and memory pressure are per-deployment runtime policy to be validated on your own stack, making these quants a concrete GPU-5/GPU-4 test candidate for the gamepc/Ollama serving endpoint against the Unsloth/GSQ baselines it already runs. Relevance holds at medium rather than high because the headline 99.63% claim was never independently replicated and attention has collapsed: this settles as a dated receipt for distrusting publisher aggregates plus a workload-specific evaluation candidate, not urgent news.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaradar:bartowski-gguf-tensor-layoutsradar:unsloth-dynamic-3-gguf-validationradar:qwen38-27b-16gb-quant-benchmarkradar:qwen-tensor-level-quant-allocationradar:qwen25-quantization-task-divergenceradar:dflash-2-parallel-drafting-validation
queries asked of Scott's wikis
  • per-tensor datatype selection quantization quality retention
  • KL divergence vs benchmark scores quantization fidelity
  • quality-speed-memory frontier local model serving
  • speculative decoding MTP draft model Vulkan Strix Halo
  • GGUF quant eval methodology BF16-normalized aggregates
  • VRAM-constrained 27B deployment notes

Measured heat

now 0 pts/hpeak 27 pts/hcomments 0/hpeers p25momentum: steady3 platformsage 625h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-15 15:27 (minted)⭐ origin echo-reconstructedByteShape's release post links this blog and downloadable ShapeLearn GGUFs, reporting 99.63% BF16-normalized aggregate benchmark performance
ByteShape on blog (echo) Β· attributed from reddit.post.1wh21e9 Β· published time unknown
β€”
09-15 14:31first on r/LocalLLaMA Β· published Β· lag ?ByteShape Qwen 3.8 27B: To KL Diverge or Not to KL Diverge, Part 2: Metric Boogaloo
enrique-byteshape
β€”
09-18 02:04first on hacker news Β· published Β· lag ?Shapelearn Qwen 3.8 27B (13.1 GB VRAM)
syntaxing
β€”
09-15 14:31amplified on r/LocalLLaMAreddit.post.1wh21e9
enrique-byteshape
peak 135 Β· 87 comments Β· 19% of case engagement
09-15 16:16amplified on r/LocalLLaMAreddit.post.1wh4v55
Thrumpwart
peak 17 Β· 4 comments Β· 2% of case engagement
09-16 20:05amplified on r/LocalLLaMAreddit.post.1wi87us
ZenZombie117
peak 1 Β· 6 comments Β· 1% of case engagement
09-18 02:04amplified on hacker newshn.story.49749393
syntaxing
peak 104 Β· 39 comments Β· 22% of case engagement
09-22 16:31amplified on r/LocalLLaMAreddit.post.1wnechu
zyxciss
peak 94 Β· 62 comments Β· 13% of case engagement
09-23 14:37amplified on r/LocalLLaMAreddit.post.1wo7sfo
Dutchnamn
peak 63 Β· 39 comments Β· 9% of case engagement
4 more amplifiers in ainews.case_chain
09-15 15:20our radar first saw it Β· lag ?discovery anchor: reddit.post.1wh21e9β€”
09-19 08:26reached heat=high Β· lag ? Β· via ledgerβ€”β€”
pace: p86 vs 1032 stories at the 336h mark (now 625h old) β€” ahead of world-labs-joins-amd (1.0x), behind rp2350-local-image-generation (1.0x)

Evidence (11) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditByteShape Qwen 3.8 27B: To KL Diverge or Not to KL Diverge, Part 2: Metric Boogaloo
LocalLLaMA
enrique-byteshape13587
🟧 echo.blog ⭐ByteShape's release post links this blog and downloadable ShapeLearn GGUFs, reporting 99.63% BF16-normalized aggregate benchmark performanceByteShapeβ€”β€”
🟠 redditByteshape releases new Qwen 3.8 27B quants and claims very impressive numbers
LocalLLaMA
Thrumpwart174
🟠 redditI retrained only the fp16 block scales of ISTA's 3-bit Qwen3.8-27B GGUF against the BF16 parent: same bytes, same loader, closer to the parent, and an honest benchmark annex. How I got there, from cutting 6% of a 30B.
LocalLLaMA
ZenZombie11716
🟧 hnShapelearn Qwen 3.8 27B (13.1 GB VRAM)syntaxing10439
🟠 redditQwen 3.8 27B at ~3 BPW on an RTX 3060: GSQ vs ByteShape IQ3-XXS 2.88BPW
LocalLLaMA
zyxciss9462
🟠 redditPerhaps the highest quality mainline quants of Qwen3.8 27B?
LocalLLaMA
Dutchnamn6239
🟠 reddit[Release & Deep Dive] ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP (i1-Q5_K_M): Sustaining 50-65+ t/s Across a FULL 128k (131,072) Context on a Single 24GB RTX 3090
LocalLLaMA
bjivanovich45
🟠 redditComparing Qwen3.8-27B fine-tunes and baselining vs. frontier
LocalLLaMA
norenEnmotalen4926
🟠 redditQwen3.8-27B: 159 tok/s on R9700, 64 tok/s on Strix Halo
LocalLLaMA
TheOriginalG217582
🟠 redditQwen3.8-27B UD-IQ4_XS Heretic + MTP on a 16 GB card with 55-68 tok/s (24gb and 12gb versions available too)
LocalLLaMA
ZestRocket3622

Interpretation history

Decision trace