2026-10-11 17:12 UTC

Redditor Training-Respect8066 claims Qwen3.8-27B at Q4_K_S with quantized context completes complex unsupervised refactors well enough that he stopped using hosted coding APIs entirely, accepting slower loops for zero marginal token cost โ€” corroborating builder substitution reports (or quality failures) would establish or refute local models displacing hosted inference for substantial coding work.

state: resolvedheat: lowuncertainty: lowconvergesscott: highlocal-models inference-economics coding-agents

What is this?

Qwen3.8-27B is Alibaba's Apache-2.0 open-weights dense coding/agentic model, released August 14, 2026 as the compact member of the Qwen3.8 family (alongside the ~2.4T-parameter Qwen3.8-Max), with 256K native context (1M hosted), an FP8 variant, tunable thinking controls, and Alibaba-reported coding scores (SWE-bench Pro 61.7) that position it as the 'worth running locally' model of its generation. The case tracks a widely-upvoted r/LocalLLaMA post whose author says the 27B at Q4_K_S with quantized KV cache handles complex unsupervised refactors well enough that he dropped hosted coding APIs entirely โ€” corroborated by several other Reddit builders across hardware tiers, one measured single-3090 run documenting a reasoning-budget failure, community benchmarks of the Flash-Next variant, and one piece of counter-evidence (a firm's rented 4xH200 DeepSeek self-hosting test roughly doubling their Claude bill โ€” a different economics regime than owned consumer hardware). Web snippets independently confirm the model's release, license, and positioning, and even corroborate the OP thread's content (Q4_K_S + Q8_0 KV, minimal tool surface, edit-tool/indentation retries as the weak point, 'top local-models post of the day'); independent economics commentary frames local as fixed-cost rather than free, with throughput and long-context penalties vs hosted batching. However, the snippets do not independently verify the total-API-abandonment claims, the Flash-Next-vs-27B benchmark table, or the reported 27B GPQA artifact (27.5% local vs ~89% published) โ€” all substitution evidence remains same-community (Reddit), and the open benchmark-replication question is unresolved in the supplied material.

Why it matters to Scott

Independent builders have now arrived in the field at the posture Scott's own stack was built around โ€” fixed-cost local inference as the working tier with hosted as escalation โ€” and the corroborated reports bear directly on live decisions in his canon: whether the LiteLLM cheap/medium/opus ladder gains a local coding tier (cheaply testable on gamepc, with the measured 3090 run as bench protocol and the reasoning-budget cap as the knob), and whether ask's suppression of native tools on the --ollama path can be lifted for a model five builders trust unsupervised. The owned-27B-vs-rented-H200 regime split and the reproducible reasoning-budget stall (connecting to the llama.cpp budget work already on the radar) hand him a sharper, qualified receipt for his local-economics writing rather than a blanket displacement claim.
dev:technology.litellmdev:concept.cost-tiered-llm-routingdev:project.gamepcdev:technology.ollamadev:concept.hardware-aware-local-inferencedev:project.askradar:qwen38-27b-local-agent-capabilityradar:concept.qwen38radar:qwen38-27b-reasoning-effortradar:mindcontrol-llamacpp-reasoning-budgetsradar:concept.local-inferenceradar:concept.inference-economicsradar:concept.coding-agents
queries asked of Scott's wikis
  • surgery problem context-sensitive modification of established code weak spot
  • local inference economics fixed cost vs metered API substitution break-even
  • quantized KV cache reasoning budget overthinking loops coding agent failure modes
  • multi-GPU serving parallel coding agents vLLM Ollama LiteLLM stack
  • coding model tier ladder when does local 27B replace hosted frontier
  • benchmark harness artifacts max_tokens answer-extraction replication controls

Measured heat

now 0 pts/hpeak 162 pts/hcomments 1/hpeers p20momentum: steady2 platformsage 264h
points/hour across evidence ยท reading as of 2026-10-05 23:42:33.461768+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-24 12:59โญ origin directly observedQwen-3.8-27B is good enough that I stopped using API
Training-Respect8066 on r/LocalLLaMA
โ€”
09-24 17:46first on r/LocalLLaMA ยท published ยท +4.8hIโ€™m moving away from cloud AI: my local-first inference setup, agent architecture, and the 3-year bet behind it
No-Name-Person111
โ€”
10-01 16:07first on hacker news ยท published ยท +171.1hFirm rents four Nvidia H200s to test '80x cheaper' DeepSeek claim
Brajeshwar
โ€”
09-24 12:59amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wp0z3i
Training-Respect8066
peak 689 ยท 287 comments ยท 71% of case engagement
09-24 17:46amplified on r/LocalLLaMAreddit.post.1wp8fh6
No-Name-Person111
peak 3 ยท 31 comments ยท 2% of case engagement
09-25 06:32amplified on r/LocalLLaMAreddit.post.1wpp28m
smallDeltaBigEffect
peak 62 ยท 32 comments ยท 7% of case engagement
09-25 20:44amplified on r/LocalLLaMAreddit.post.1wq7e74
AdInternational5848
peak 75 ยท 51 comments ยท 9% of case engagement
09-26 06:54amplified on r/LocalLLaMAreddit.post.1wqjr3f
jacek2023
peak 22 ยท 23 comments ยท 3% of case engagement
10-01 16:07amplified on hacker newshn.story.49923539
Brajeshwar
peak 2 ยท 0 comments ยท 0% of case engagement
4 more amplifiers in ainews.case_chain
09-24 13:20our radar first saw it ยท +0.3hdiscovery anchor: reddit.post.1wp0z3iโ€”

Evidence (10) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญQwen-3.8-27B is good enough that I stopped using API
LocalLLaMA
Training-Respect8066680287
๐ŸŸ  redditIโ€™m moving away from cloud AI: my local-first inference setup, agent architecture, and the 3-year bet behind it
LocalLLaMA
No-Name-Person111031
๐ŸŸ  redditDid anyone do a full bench of e.g. Qwen Flash Next IQ4 and Qwen 27b FP8? Here are some
LocalLLaMA
smallDeltaBigEffect6232
๐ŸŸ  reddit4-5 days replacing Claude w Qwen 3.8 Next
LocalLLaMA
AdInternational58487551
๐ŸŸ  redditQwen3.8-27B Q4_K_M on 2x3060
LocalLLaMA
jacek20232223
๐ŸŸง hnFirm rents four Nvidia H200s to test '80x cheaper' DeepSeek claimBrajeshwar20
๐ŸŸ  redditQwen3.8-27B Q4_K_M on one RTX 3090 + OpenCode: throughput, four coding tasks, and a reasoning-budget failure
LocalLLaMA
Adorable-Cost-3249213
๐ŸŸ  redditTuned/abliterated Qwen3.8-27b into a 24gb card 262k guff using the newest unreleased version of LexiPanel. It's fast with reliable draft acceptance. Made for 7900xtx but should work on whatever 24gb card with this setup and headless. Doesn't get dumber while coding like most of the other fine-tunes.
LocalLLaMA
W61k3r625
๐ŸŸ  reddit2x 3090 in server chassis, upgrade time
LocalLLaMA
knighty1981014
๐ŸŸ  redditMoving from Qwen 27B to cloud agents was eye-opening. But I have no regrets.
LocalLLaMA
GrungeWerX038

Interpretation history

Decision trace