2026-10-11 16:38 UTC

TacGibs releases an 8-bit AWQ+GPTQ hybrid quantization of Qwen 3.8 27B Swift 1.5 achieving ~100 tok/s decode on dual RTX 3090 with vLLM+MTP while reducing overthinking versus the base model.

state: seedheat: mediumuncertainty: mediumconvergesscott: highqwen38-quantization awq-gptq-hybrid dual-3090-inferenceTacGibs

What is this?

Community quantizer TacGibs (Reddit user) published an 8-bit AWQ+GPTQ hybrid quantization of Qwen 3.8 27B "Swift 1.5" targeting dual RTX 3090 (48 GB VRAM) with vLLM and MTP speculative decoding, claiming ~100 tok/s decode and reduced overthinking versus the base model. The Swift naming appears to be a community variant lineage (distinct from Qwen's official "Flash-Next") built on the same Qwen 3.8 27B base. Corroborating signals: syv-ai/HyperQwen achieves 127 tok/s single-stream on one 3090 with vLLM patches and requantization; Andrew Zhu reports 177 tok/s on one 3090 Ti; Hermes AI cites 417 tok/s batched on one 3090. The dual-3090 8-bit claim is consistent with the VRAM headroom (27B ร— 8-bit โ‰ˆ 27 GB + KV + MTP overhead fits in 48 GB). Snippets do not directly confirm the AWQ+GPTQ hybrid method or the "overthinking reduction" claim โ€” those rest on the Reddit post alone.

Why it matters to Scott

A community quantizer independently produces a hybrid AWQ+GPTQ 8-bit artifact for dual RTX 3090 (48 GB) with vLLM+MTP speculative decoding โ€” exactly the hardware-aware local-inference pattern Scott's frameworks treat as sovereign infrastructure. The claimed overthinking reduction directly engages his reasoning-paradox/over-agreeableness work and fast-slow split architecture. This is not merely an example; it is a new dated artifact with testable speed/quality/reasoning tradeoffs that would change what he deploys on gamepc and argues in sovereignty positions.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.model-perishabilityip:framework.sovereign-software-assuranceip:concept.ai-over-agreeablenessip:concept.reasoning-paradoxip:framework.fast-slow-splitdev:concept.metacognitive-resolution-controlradar:bartowski-gguf-tensor-layoutsradar:convrot-llama-cpp-quantizationradar:adaptive-speculative-decoding-300-gpuradar:consumer-gpu-p2p-kernel-forkradar:concept.open-weightsradar:geometry-preserving-nvfp4-distillation
queries asked of Scott's wikis
  • local-inference quantization AWQ GPTQ hybrid 8-bit
  • dual-3090 48GB VRAM serving vLLM MTP speculative-decoding
  • qwen3-8-27b quantization artifacts Swift HyperQwen
  • overthinking reduction reasoning-models quantization tradeoffs
  • consumer-GPU sovereignty local-serve open-weights strategy

Measured heat

now 0 pts/hpeak 2 pts/hcomments 0/hpeers p18momentum: steady1 platformsage 24h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-10 16:11โญ origin directly observed48Gb VRAM speed AND quality ! (Qwen 3.8 27B Swift 1.5 W8A16)
TacGibs on r/LocalLLaMA
โ€”
10-10 16:11amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1x2j3gx
TacGibs
peak 7 ยท 8 comments ยท 100% of case engagement
10-10 18:35our radar first saw it ยท +2.4hdiscovery anchor: reddit.post.1x2j3gxโ€”
pace: p65 vs 923 stories at the 12h mark (now 24h old) โ€” ahead of anthropic-usage-policy-election-interference-ban (1.1x), behind bfl-flux-3-image-release (0.9x)

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญ48Gb VRAM speed AND quality ! (Qwen 3.8 27B Swift 1.5 W8A16)
LocalLLaMA
TacGibs78

Interpretation history

Decision trace