Qwen3.8-27B is Alibaba's Apache-2.0 open-weights dense coding/agentic model, released August 14, 2026 as the compact member of the Qwen3.8 family (alongside the ~2.4T-parameter Qwen3.8-Max), with 256K native context (1M hosted), an FP8 variant, tunable thinking controls, and Alibaba-reported coding scores (SWE-bench Pro 61.7) that position it as the 'worth running locally' model of its generation. The case tracks a widely-upvoted r/LocalLLaMA post whose author says the 27B at Q4_K_S with quantized KV cache handles complex unsupervised refactors well enough that he dropped hosted coding APIs entirely โ corroborated by several other Reddit builders across hardware tiers, one measured single-3090 run documenting a reasoning-budget failure, community benchmarks of the Flash-Next variant, and one piece of counter-evidence (a firm's rented 4xH200 DeepSeek self-hosting test roughly doubling their Claude bill โ a different economics regime than owned consumer hardware). Web snippets independently confirm the model's release, license, and positioning, and even corroborate the OP thread's content (Q4_K_S + Q8_0 KV, minimal tool surface, edit-tool/indentation retries as the weak point, 'top local-models post of the day'); independent economics commentary frames local as fixed-cost rather than free, with throughput and long-context penalties vs hosted batching. However, the snippets do not independently verify the total-API-abandonment claims, the Flash-Next-vs-27B benchmark table, or the reported 27B GPQA artifact (27.5% local vs ~89% published) โ all substitution evidence remains same-community (Reddit), and the open benchmark-replication question is unresolved in the supplied material.
Independent builders have now arrived in the field at the posture Scott's own stack was built around โ fixed-cost local inference as the working tier with hosted as escalation โ and the corroborated reports bear directly on live decisions in his canon: whether the LiteLLM cheap/medium/opus ladder gains a local coding tier (cheaply testable on gamepc, with the measured 3090 run as bench protocol and the reasoning-budget cap as the knob), and whether ask's suppression of native tools on the --ollama path can be lifted for a model five builders trust unsupervised. The owned-27B-vs-rented-H200 regime split and the reproducible reasoning-budget stall (connecting to the llama.cpp budget work already on the radar) hand him a sharper, qualified receipt for his local-economics writing rather than a blanket displacement claim.
dev:technology.litellmdev:concept.cost-tiered-llm-routingdev:project.gamepcdev:technology.ollamadev:concept.hardware-aware-local-inferencedev:project.askradar:qwen38-27b-local-agent-capabilityradar:concept.qwen38radar:qwen38-27b-reasoning-effortradar:mindcontrol-llamacpp-reasoning-budgetsradar:concept.local-inferenceradar:concept.inference-economicsradar:concept.coding-agents
queries asked of Scott's wikis
- surgery problem context-sensitive modification of established code weak spot
- local inference economics fixed cost vs metered API substitution break-even
- quantized KV cache reasoning budget overthinking loops coding agent failure modes
- multi-GPU serving parallel coding agents vLLM Ollama LiteLLM stack
- coding model tier ladder when does local 27B replace hosted frontier
- benchmark harness artifacts max_tokens answer-extraction replication controls
2026-10-05T12:51:23Z
rmhubbert's comment โ a second independent claim of six months' exclusive local coding, posted as a direct rebuttal to the defector โ patches the last fragility this discussion cycle could address (total-API-abandonment resting on the OP alone). With every object at rate floor, the counter-post dead on arrival, and no new evidence kind in flight, the episode's finding is finalized rather than extended: the hypothesis's own establishment condition (corroborating builder substitution reports) has been met in qualified form. Resolve absorbed; a harness-controlled replication, cross-platform pickup, or the named Qwen 4 Flash trigger would open the successor episode.
2026-10-04T22:27:29Z
First dedicated reverse-migration counter-report arrives: GrungeWerX, a self-identified long-time Qwen-local builder, moved to cloud agents and found the quality gap 'eye-opening' โ the counter-side now has a practitioner witness with claimed receipts, beyond inline doubters and the H200 economics story, sharpening the regime boundary rather than overturning the pattern (his quants/hardware are unanswered and the community rejected the post at 0.13 ratio). Hold corroborated; heat stays low despite the magnitude-valve flag, which prices the Sept 24-26 burst โ everything is now at rate floor and the counter-post died on arrival.
2026-10-04T22:25:54Z
evidence attached: reddit.post.1wxrvpj โ Counter-evidence: a committed Qwen-local builder moved to cloud agents and found the quality gap eye-opening.
2026-10-03T05:31:41Z
Third consecutive no-advancement look: the newest attachment (knighty1981's dual-3090 hardware-upgrade thread, 0pts, 0.5 ratio) is the weakest evidence class yet โ incidental mention of running 27B/256k via OpenCode for admin tasks inside a chassis-upgrade discussion, not a displacement report โ and comment churn on the abliterated-artifact post still leaves its quality question unanswered. Meaning unchanged: a Reddit-corroborated, regime-qualified substitution pattern in a holding pattern until a new KIND of evidence arrives (cross-platform pickup, harness-controlled replication, large-codebase durability, cross-community artifacts, or the named Qwen 4 Flash trigger).
2026-10-02T23:24:08Z
evidence attached: reddit.post.1ww7rxb โ Independent builder running Qwen3.8-27B at 256K context on dual 3090s via OpenCode for sustained real workloads โ another local-substitution data point the case should carry.
2026-10-02T14:47:37Z
First distribution-layer artifact arrives: W61k3r's HF upload packages 27B plus ~262k context onto one 24GB card (LexiPanel fit planner + MTP draft), a sign the ecosystem is starting to productize the pattern rather than just report it โ but it is an abliterated derivative with self-reported, unverified quality (unanswered 'how did you measure quality?', KV-quant degradation reports in comments), same-community and thin, and matches none of the case's named advancement triggers. Consolidation, not advancement: hold corroborated/low.
2026-10-02T14:26:17Z
evidence attached: reddit.post.1wvtxkx โ Another builder running daily agentic coding on local Qwen3.8-27B with 110โ170k context at 36โ41 tok/s on one 24GB card, independent support for the local-substitution hypothesis.
2026-10-01T22:28:01Z
grounded: converges/high โ Independent builders have now arrived in the field at the posture Scott's own stack was built around โ fixed-cost local inference as the working tier with hoste
2026-10-01T22:18:35Z
The new item is the first measured, failure-inclusive implementation writeup at the single-24GB-GPU tier (Adorable-Cost-3249: RTX 3090 + OpenCode + llama.cpp, throughput across four coding tasks plus a concrete reasoning-budget stall), upgrading the GPU-poor picture from comment-level claims (tsangberg, ailee43) to a documented run and confirming the overthinking-loop failure as the pattern's reproducible edge with a tunable mitigation. But it is still same-community and thin (3pts), so by the case's own bar it consolidates rather than advances: hold corroborated/low, still waiting on a new KIND of evidence โ cross-platform pickup, harness-controlled replication, large-codebase durability, or the named external trigger (Qwen 4 Flash).
2026-10-01T20:35:47Z
evidence attached: reddit.post.1wv9rzc โ Another builder's measured real-world Qwen3.8-27B local coding-agent run โ including a concrete reasoning-budget failure โ bears directly on whether local models displace hosted coding APIs.
2026-10-01T17:39:26Z
First cross-platform evidence arrives as counter-evidence: a firm's 4xH200 rental test found self-hosting DeepSeek roughly doubles the Claude bill โ it doesn't touch the owned-hardware 27B substitution reports (rented datacenter large-MoE vs consumer dense 27B are different economics regimes) but forces the displacement claim to be regime-qualified rather than blanket. Corroboration holds at five builders, all still Reddit-only, with no new positive periphery in ~6 days; the case stays corroborated and cools.
2026-10-01T16:32:01Z
evidence attached: hn.story.49923539 โ Independent H200 rental test finding self-hosting DeepSeek doubles the Claude bill is direct counter-evidence on open models displacing hosted inference for coding.
2026-09-26T07:26:49Z
Fifth independent builder (jacek2023, 4x3090) runs Qwen3.8-27B as his daily multi-agent setup with parallel=2 โ the first report of parallel agents offsetting the local speed penalty, and dependence deep enough that repossessing the GPU box for vLLM left him 'missing my AI' โ but it is the same community and same pattern at thin engagement (5pts/1 comment), and the main thread has cooled to ~3 pts/h from a ~135 peak. This consolidates corroboration without changing the case's shape: hold at corroborated, cool to low; the next real move requires a new KIND of evidence (cross-platform pickup, harness-controlled benchmark replication, large-codebase durability), not more same-community anecdotes.
2026-09-26T07:23:35Z
evidence attached: reddit.post.1wqjr3f โ Builder running two parallel coding agents on local Qwen3.8-27B across machines as his daily setup โ another independent substitution-pattern data point the case explicitly seeks.
2026-09-25T21:37:10Z
Fourth independent builder (AdInternational5848) replaces hosted Claude with local Qwen 3.8 Next on an M1 Ultra for 4-5 days โ the first multi-day durability report โ though on the Flash-Next variant with high-end unified memory rather than the 27B on modest hardware, so corroboration consolidates (4 builders + benchmark line) while staying Reddit-only and cooling: hold at corroborated, don't promote. The benchmark thread's velocity spike is substantive reproducibility work, not amplification โ commenters suspect the 27B's GPQA 27.5% is an answer-extraction/max_tokens artifact, an open caveat that gates any independent verification.
2026-09-25T21:25:28Z
evidence attached: reddit.post.1wq7e74 โ Another builder independently reports replacing hosted Claude with local Qwen 3.8 on an M1 Ultra for real coding work โ corroborating evidence for the local-substitution hypothesis.
2026-09-25T07:24:14Z
A third evidence line arrives: smallDeltaBigEffect's independent benchmarks show Flash-Next IQ4_XS clearly outscores 27B FP8 on reasoning/knowledge (GPQA 42.5 vs 27.5) while 27B wins decisively on practicality (12GB at Q3, ~3.5x faster), and an in-thread quant-vs-published-score gap adds a reproducibility caveat; meanwhile the second corroborating post (No-Name-Person111) has collapsed to 0.43 ratio amid AI-slop accusations with a moderator query, so corroboration now rests on the in-thread builders plus the benchmark line, not that post. Engagement peaked (~95 pts/h) and is cooling on a single platform โ corroborated but not accelerating; drop to low absent durability reports or cross-platform pickup.
2026-09-25T07:22:46Z
evidence attached: reddit.post.1wpp28m โ Independent quality/speed benchmarks of Qwen3.8 27B FP8 vs Flash-Next quants bear directly on whether these local models are capable enough to displace hosted coding APIs.
2026-09-24T19:20:07Z
The thread's own comments now carry independent builder confirmations on distinct stacks (nvfp4-mtp + Cline in LM Studio; IQ4_XS on 16GB as sole dev/cybersec model), and a second, separate local-first migration post corroborates the direction โ lifting this from a lone anecdote to a corroborated substitution pattern. Credible tempering (10h vs 30min, 'smaller projects only') keeps total-API-abandonment qualified; measured heat (98th percentile, single platform, steady) outrates the 'low' label but doesn't justify high.
2026-09-24T18:29:04Z
evidence attached: reddit.post.1wp8fh6 โ Second independent builder report of deliberately abandoning hosted inference for a local-first stack with cloud only on demand, directly corroborating the substitution case.
2026-09-24T13:38:06Z
grounded: converges/high โ A firsthand total-substitution report โ hosted coding APIs abandoned entirely for a Q4_K_S local 27B doing complex unsupervised refactors โ independently arrive
2026-09-24T13:30:20Z
case created โ Distinct firsthand substitution claim โ no open case covers a builder dropping hosted APIs for local Qwen3.8-27B refactoring work.