2026-10-11 17:12 UTC

Storterald claims IQ4_XS variants offer the best coding-quality tradeoff among 21 tested Qwen3.8 27B quantizations that fit on a 16GB RTX 5080, informing practical deployment choices for memory-constrained local inference.

state: expiredheat: lowuncertainty: highknownscott: mediumqwen38 quantization local-inference gpu-benchmarksStorteraldQwen

What is this?

The case concerns a reported comparison of 21 quantized variants of Qwen3.8 27B for local coding workloads on a 16GB RTX 5080, attributed to Storterald, although the supplied snippets do not directly identify or substantiate that author or the full benchmark. The surrounding evidence confirms that IQ4_XS and hybrid IQ4_XS/IQ3_S builds are designed to fit this model within a strict 16GB VRAM budget using llama.cpp, with context length competing for the same limited memory. Support for IQ4_XS as the uniquely best quality–performance tradeoff is thin or mixed: one account found it only slightly better, while another benchmark favored a roughly 17GB Q4_K_M build when it fits.

Why it matters to Scott

The radar already tracks this model’s local-agent performance under “qwen38-27b-local-agent-capability,” while Scott’s “gamepc — self-hosted GPU model zoo” and “Hardware-aware local inference” make quantization-versus-VRAM evidence potentially actionable for model selection. The claimed IQ4_XS result could affect deployment choices on comparable constrained hardware, but the thin and mixed benchmark support limits its present decision value.
dev:project.gamepcdev:concept.hardware-aware-local-inferenceradar:qwen38-27b-local-agent-capabilityradar:concept.quantizationradar:concept.local-inference
queries asked of Scott's wikis
  • local-model quantization quality versus VRAM tradeoffs
  • 16GB GPU local coding-agent deployment
  • llama.cpp quantization and context-memory budgeting
  • local inference benchmark methodology for coding models
  • consumer GPU economics versus hosted model APIs
  • hybrid per-layer quantization for agent workloads

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐I benchmarked 21 Qwen3.8 27B variants on 16GB VRAM
LocalLLaMA
Storterald28383
🟠 redditVillager Simulation Game POC Created with Qwen3.8-27B-UD-Q3_K_XL.gguf - 16GB VRAM
LocalLLaMA
Fancy-Snow76337
🟠 redditQwen 3.8-27B NVFP4 actually beats Q5_K_M and touches official BF16 levels, but only if you change 2 sampling params. Also almost 3x faster.
LocalLLaMA
UmpireBorn3719030

Interpretation history

Decision trace