Qwen3.8-27B is an open-weight 27B model from Alibaba's Qwen team (August 2026) that ships with an official `reasoning_effort` knob — `xhigh` (default), `medium`, `low` — and multiple independent sources confirm the default is the most expensive setting. Simon Willison documented 'spectacular overthinking' on consumer hardware (a pelican-SVG prompt took 21 minutes and 22k+ reasoning tokens), VentureBeat relays Artificial Analysis figures showing roughly 4x the output-token volume of comparable open-weight models, and Qwen's own tracker carries the operational fallout: the chat template rejects Claude Code's `high` value with a vLLM HTTP 500, and issue #216 reports `xhigh` returning empty answers on ~17% of calls. Independent benchmarking (kodesage) shows a real tradeoff — `xhigh` scores highest but is slowest, `medium` is fastest and near-par in quality, `low` underperforms — and local backends handle the knob inconsistently (LM Studio ignores it, Ollama lowers it silently). The snippets strongly corroborate the xhigh default as a material cost/latency hazard, but they do not cover the case's original reporter DerTomsn or his M5 Max measurement, which remains same-source testimony, and they do not establish that explicit effort selection lowers cost per successful task once retries, truncation and failures are counted.
Scott's own canon already carries this position: ip:concept.high-not-max argues precisely that reasoning effort is a per-call budget and max-effort defaults are a trap, and his 12-factor/production pages treat silent defaults in local stacks as a first-class hazard — the accumulated evidence (OMP/vLLM deployment corroboration, the 72-comment loop thread, the Swift post-training detour) re-derives the synthesis his context-engineering and AI-unit-economics pages already hold: the effort knob is necessary but insufficient, and cost-per-successful-task including retries and truncation is the real metric. It stays medium rather than low because it is a concrete audit trigger, not just another example: backends demonstrably mishandle the knob (LM Studio ignores it, Ollama silently lowers it, vLLM's template 500s on 'high', xhigh returns empty answers ~17% per Qwen's tracker), so effort propagation and accepted values in ask's local OpenAI-compatible path and the LiteLLM gateway deserve a check on his active stack.
ip:concept.high-not-maxip:framework.12-factor-agents-frameworkip:concept.ai-unit-economicsip:framework.context-engineeringdev:project.askdev:technology.litellmradar:qwen38-27b-reasoning-effortradar:ukisai-swift-family-releaseradar:claude-code-effort-controlsradar:mindcontrol-llamacpp-reasoning-budgetsradar:ollama-silent-context-truncationradar:frontierharness-17x-cost-variationradar:concept.reasoning-tokens
queries asked of Scott's wikis
- reasoning effort parameter propagation in my OpenAI-compatible harness code
- 'High, Not Max' — my position on effort/settings headroom
- cost per successful task: token accounting with retries and truncation
- silent-default configuration hazards in local model stacks (LM Studio/Ollama/llama.cpp/vLLM)
- thinking-token budgets and TTFT/latency ceilings in agent loops
- config vs post-training levers for cutting reasoning tokens
2026-10-08T04:35:37Z
A transient velocity spike on the real-user thread (1wyp4n2) added ~34 comments discussing thinking loops and context exhaustion — the practical failure mode the case already tracks — but measured heat has since returned to 0 pts/h (steady, 50th percentile). The Swift-wave expansion that drove high heat in mid-September has fully crested; no new implementations, communities, or outlets have appeared. The case's meaning remains settled: the silent xhigh default is an established operational hazard worth auditing, but effort selection alone does not bound cost; loop behavior and context exhaustion dominate real workloads.
2026-10-06T04:23:51Z
grounded: known/medium — Scott's own canon already carries this position: ip:concept.high-not-max argues precisely that reasoning effort is a per-call budget and max-effort defaults are
2026-10-06T04:15:21Z
The Swift-wave expansion that earned high heat on 09-19 has crested: no new implementations, communities or outlets since, and current activity is one live contested real-user thread on thinking loops exhausting the 95k context — so attention cools to medium. The case's meaning has settled: the silent xhigh default is an established operational hazard worth auditing, but accumulating evidence shows effort selection alone does not bound cost; loop behavior and context exhaustion dominate real workloads.
2026-10-06T03:33:52Z
evidence attached: reddit.post.1wyp4n2 — 72-comment real-user report of thinking loops consuming the full 95k context is the practical failure mode the reasoning-effort-default case tracks.
2026-09-19T08:27:35Z
The efficiency episode now warrants high attention because Swift adoption, runnable derivative packaging and cross-platform discussion show expanding reach, not because the default-effort savings claim has been proved. The latest discussion adds no controlled result: auditing effort propagation remains actionable, while cost per successful coding task is unresolved.
2026-09-19T02:30:23Z
An independent OMP/vLLM coding-agent deployment reports that unset roles defaulted to xhigh, corroborating the operational configuration hazard beyond the original comparison. Its reported latency improvement followed several simultaneous changes, so it supports auditing the whole harness rather than attributing savings to effort selection alone.
2026-09-19T02:21:36Z
evidence attached: reddit.post.1wk8jef — Independent coding-agent use confirms that explicit effort settings, output limits, and context handling materially affect Qwen3.8 workflow latency and cost.
2026-09-17T20:45:19Z
The latest Swift post is an adoption milestone for the existing derivative, not a new release or efficiency result; a reported NVFP4-to-GGUF conversion adds packaging context but no demonstrated coding-workload benefit. Neither changes the unresolved question of whether explicitly lowering base-model effort reduces cost per successful task.
2026-09-17T20:23:00Z
evidence attached: reddit.post.1wj3s31 — The release provides adoption evidence and a claimed 58.3% token reduction for a Qwen 3.8 27B finetune, materially informing the open model's practical cost and throughput tradeoffs.
2026-09-16T15:40:00Z
The HN attachment repeats Swift's creator claims rather than independently validating them; its discussion reinforces the known risk that loops persist at lower effort. Cross-platform coverage does not strengthen the evidence that explicit base-model effort selection lowers cost per successful coding task.
2026-09-16T15:22:38Z
evidence attached: hn.story.49727511 — shared external link with case evidence
2026-09-15T18:00:32Z
The new post reports independent benchmarking of the existing Swift derivative, not a separate ThinkingCap fine-tune; its roughly 40% reasoning-token reduction adds qualified support for post-training efficiency gains. This still does not validate savings from base-model effort selection or lower cost per successful coding task, and the creator explicitly disputes training on ThinkingCap traces.
2026-09-15T17:26:08Z
evidence attached: reddit.post.1wh5elt — Community fine-tune cutting the same model's reasoning tokens ~40% is material context for whether Qwen3.8-27B reasoning-cost management is a real lever.
2026-09-15T07:31:53Z
The latest discussion adds endorsements but no new controlled results beyond the already-accounted-for Swift speed/completeness tradeoff. Explicit effort configuration remains actionable, but neither growing attention nor derivative-model enthusiasm establishes lower cost per successful coding task.
2026-09-14T23:29:04Z
A limited third-party Swift test now reports faster completion but omitted details, adding a concrete deployment result that qualifies the creators' near-lossless efficiency claim without disproving their benchmark results. This strengthens the case for measuring completeness alongside speed, not for assuming lower cost per successful coding task.
2026-09-14T16:33:06Z
Swift-Qwen3.8-27B adds a concrete post-training intervention with creator-reported efficiency gains, moving this beyond configuration anecdotes to a candidate implementation worth testing. It does not establish that selecting lower effort on the base model delivers comparable savings, or that the derivative beats base medium on cost per successful coding task.
2026-09-14T16:22:49Z
evidence attached: reddit.post.1wg7dd5 — The released Swift-Qwen3.8-27B artifact directly tests whether reducing excessive reasoning tokens can preserve quality while materially lowering local inference cost.
2026-09-13T10:24:52Z
The new discussion adds conflicting effort-specific anecdotes: medium reportedly works well for one user, while another reports persistent looping even on low. This reinforces workload-dependent tuning rather than establishing that lowering the default reduces cost per successful task.
2026-09-13T10:21:43Z
evidence attached: reddit.post.1wf39vc — User discussion directly bears on whether Qwen3.8's reasoning defaults create material quality and cost differences on simple tasks.
2026-09-10T21:41:05Z
The refreshed discussion remains amplification and coding-use testimony, not independent evidence of the effort setting's cost impact. Explicit effort selection remains a useful configuration check, but lower cost per successful coding task is still unestablished.
2026-09-09T21:30:11Z
The refreshed discussion remains coding-use testimony rather than independent evidence about effort settings; it does not strengthen the claimed cost benefit. Explicit effort propagation remains a useful configuration check, but savings per successful task—including retries and quality tradeoffs—are still unestablished.
2026-09-09T20:39:05Z
The refreshed discussion adds coding-use anecdotes and reports of looping, but no independent effort-controlled replication; the blog echo remains the same author's testimony. The documented xhigh default warrants a configuration check, while lower effort's effect on cost per successful task remains unestablished.
2026-09-09T20:28:59Z
grounded: known/medium — The radar already tracks this model’s medium-versus-xhigh coding tradeoff in radar:qwen38-27b-reasoning-effort; the documented default adds an actionable config
2026-09-09T20:24:19Z
case created — A quoted implementation detail and controlled local measurement establish a distinct, actionable inference-cost episode, although the supplied excerpt does not establish the quality tradeoff.