2026-10-11 16:38 UTC

DerTomsn reports that Qwen3.8-27B silently defaults to its most expensive xhigh reasoning setting through its chat template, making explicit effort selection a potentially material latency and compute-cost control for local coding workloads.

state: corroboratedheat: mediumuncertainty: mediumknownscott: mediumlocal-inference inference-economics coding-agentsDerTomsnQwen
Surfaced 2026-09-19T08:27:35Z — priced heat=high at reprice: The efficiency episode now warrants high attention because Swift adoption, runnable derivative packaging and cross-platform discussion show expanding reach, not because the default-effort savings claim has been proved. The latest discussion adds no controlled result: auditing effort propagation remains actionable, while cost per successful coding task is unresolved.

What is this?

Qwen3.8-27B is an open-weight 27B model from Alibaba's Qwen team (August 2026) that ships with an official `reasoning_effort` knob — `xhigh` (default), `medium`, `low` — and multiple independent sources confirm the default is the most expensive setting. Simon Willison documented 'spectacular overthinking' on consumer hardware (a pelican-SVG prompt took 21 minutes and 22k+ reasoning tokens), VentureBeat relays Artificial Analysis figures showing roughly 4x the output-token volume of comparable open-weight models, and Qwen's own tracker carries the operational fallout: the chat template rejects Claude Code's `high` value with a vLLM HTTP 500, and issue #216 reports `xhigh` returning empty answers on ~17% of calls. Independent benchmarking (kodesage) shows a real tradeoff — `xhigh` scores highest but is slowest, `medium` is fastest and near-par in quality, `low` underperforms — and local backends handle the knob inconsistently (LM Studio ignores it, Ollama lowers it silently). The snippets strongly corroborate the xhigh default as a material cost/latency hazard, but they do not cover the case's original reporter DerTomsn or his M5 Max measurement, which remains same-source testimony, and they do not establish that explicit effort selection lowers cost per successful task once retries, truncation and failures are counted.

Why it matters to Scott

Scott's own canon already carries this position: ip:concept.high-not-max argues precisely that reasoning effort is a per-call budget and max-effort defaults are a trap, and his 12-factor/production pages treat silent defaults in local stacks as a first-class hazard — the accumulated evidence (OMP/vLLM deployment corroboration, the 72-comment loop thread, the Swift post-training detour) re-derives the synthesis his context-engineering and AI-unit-economics pages already hold: the effort knob is necessary but insufficient, and cost-per-successful-task including retries and truncation is the real metric. It stays medium rather than low because it is a concrete audit trigger, not just another example: backends demonstrably mishandle the knob (LM Studio ignores it, Ollama silently lowers it, vLLM's template 500s on 'high', xhigh returns empty answers ~17% per Qwen's tracker), so effort propagation and accepted values in ask's local OpenAI-compatible path and the LiteLLM gateway deserve a check on his active stack.
ip:concept.high-not-maxip:framework.12-factor-agents-frameworkip:concept.ai-unit-economicsip:framework.context-engineeringdev:project.askdev:technology.litellmradar:qwen38-27b-reasoning-effortradar:ukisai-swift-family-releaseradar:claude-code-effort-controlsradar:mindcontrol-llamacpp-reasoning-budgetsradar:ollama-silent-context-truncationradar:frontierharness-17x-cost-variationradar:concept.reasoning-tokens
queries asked of Scott's wikis
  • reasoning effort parameter propagation in my OpenAI-compatible harness code
  • 'High, Not Max' — my position on effort/settings headroom
  • cost per successful task: token accounting with retries and truncation
  • silent-default configuration hazards in local model stacks (LM Studio/Ollama/llama.cpp/vLLM)
  • thinking-token budgets and TTFT/latency ceilings in agent loops
  • config vs post-training levers for cutting reasoning tokens

Measured heat

now 0 pts/hpeak 10 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 764h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-09 20:24 (minted)⭐ origin echo-reconstructedThe author quotes a chat-template default of reasoning_effort='xhigh' and reports comparing low, medium, and xhigh on one M5 Max with the sa
DerTomsn on blog (echo) · attributed from reddit.post.1wbwnlx · published time unknown
—
09-09 20:05first on r/LocalLLaMA · published · lag ?Qwen3.8-27B has the best coding ceiling you can run at home on consumer hardware, it ships with reasoning_effort defaulting to xhigh - I measured what that costs
DerTomsn
—
09-16 14:24first on hacker news · published · lag ?Show HN: Swift-Qwen3.8-27B, -58.3% thinking, x1.95 speed, accuracy of xhigh
kisjovan
—
09-09 20:05amplified on r/LocalLLaMAreddit.post.1wbwnlx
DerTomsn
peak 65 · 39 comments · 3% of case engagement
09-13 09:43amplified on r/LocalLLaMAreddit.post.1wf39vc
thoquz
peak 10 · 46 comments · 1% of case engagement
09-14 15:57amplified on r/LocalLLaMAreddit.post.1wg7dd5
Secure_Recording_472
peak 950 · 402 comments · 35% of case engagement
09-15 16:36amplified on r/LocalLLaMAreddit.post.1wh5elt
returnity
peak 214 · 90 comments · 8% of case engagement
09-16 14:24amplified on hacker newshn.story.49727511
kisjovan
peak 33 · 17 comments · 2% of case engagement
09-17 19:30amplified on r/LocalLLaMA 👑reddit.post.1wj3s31
Secure_Recording_472
peak 990 · 618 comments · 42% of case engagement
2 more amplifiers in ainews.case_chain
09-09 20:20our radar first saw it · lag ?discovery anchor: reddit.post.1wbwnlx—
09-19 08:27reached heat=high · lag ? · via ledger——
pace: p96 vs 519 stories at the 720h mark (now 764h old) — ahead of amodei-frontier-pacing-commitment (1.0x), behind meta-muse-spark-13-release (1.0x)

Evidence (9) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditQwen3.8-27B has the best coding ceiling you can run at home on consumer hardware, it ships with reasoning_effort defaulting to xhigh - I measured what that costs
LocalLLaMA
DerTomsn6229
🟧 echo.blog ⭐The author quotes a chat-template default of reasoning_effort='xhigh' and reports comparing low, medium, and xhigh on one M5 Max with the saDerTomsn——
🟠 redditHow does Qwen 3.8 27B compare on low thinking mode to the older 3.6 models?
LocalLLaMA
thoquz946
🟠 redditUkisAI Swift-Qwen3.8-27B / -58.3% thinking, x1.95 speed while keeping the accuracy of xhigh
LocalLLaMA
Secure_Recording_472950398
🟠 redditCut Qwen3.8-27B Reasoning Tokens by 40% -- 3.8 'ThinkingCap' benchmarked!
LocalLLaMA
returnity21490
🟧 hnShow HN: Swift-Qwen3.8-27B, -58.3% thinking, x1.95 speed, accuracy of xhighkisjovan3317
🟠 redditThank you :) Swift Qwen 3.8 27B now has 100k+ downloads, is #1 finetune and #9 model on HuggingFace Trending
LocalLLaMA
Secure_Recording_472989617
🟠 redditTuning Qwen 3.8 27B and OMP as a coding agent on 2× 3090s
LocalLLaMA
bolts983913
🟠 redditQwen 3.8 27b just feels… ok?
LocalLLaMA
YeetHub85203

Interpretation history

Decision trace