2026-10-11 17:12 UTC

Independent benchmarks will determine whether Qwen3.8-27B’s medium reasoning mode offers a better agentic-coding quality and token-efficiency tradeoff than xhigh mode and Qwen3.6.

state: resolvedheat: lowuncertainty: mediumknownscott: highlocal-inference open-models coding-agents inference-economicsAlibabaQwen

What is this?

The case concerns an independent local agentic-coding benchmark of Alibaba’s Qwen3.8-27B, comparing model weights, quantizations, inference engines, cache quantization, and reasoning-effort settings, with Qwen3.6 as a baseline. One snippet says β€œmedium” applies no extra reasoning directive while β€œxhigh” adds instructions to think carefully and validate assumptions, making quality versus token use the central tradeoff. However, the supplied snippets provide no direct comparative scores establishing that medium beats xhigh or Qwen3.6; they offer only broader Qwen benchmark context, so the hypothesis remains unverified here.

Why it matters to Scott

Scott already holds the central position in β€œHigh, Not Max”: sustained agents should avoid maximum per-call reasoning unless its added value is demonstrated. This benchmark could directly refine reasoning defaults for his local coding-agent stack and hardware-aware inference work, but the supplied evidence contains no comparative results yet, so it does not currently validate, challenge, or extend that position.
ip:concept.high-not-maxip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisondev:project.gamepcdev:concept.hardware-aware-local-inferenceradar:claude-code-effort-controlsradar:mindcontrol-llamacpp-reasoning-budgetsradar:tokenspeed-qwen38-serving-validation
queries asked of Scott's wikis
  • reasoning effort versus token efficiency in coding agents
  • local coding-model benchmark methodology
  • quantization and KV-cache effects on agent reliability
  • open-model inference economics for coding agents
  • adaptive reasoning budgets in agent harnesses
  • local inference engine comparisons for agentic workloads

Measured heat

no measured readings yet β€” the hourly heat pass fills this in

How the heat travelled

no chain yet β€” the hourly chain pass fills this in

Evidence (59) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Local agentic coding Benchmark : Qwen 3.8 27B (in many weights quants / cache quants / engine / reasoning effort) vs others.
LocalLLaMA
WonderRico4617
🟠 redditQwen 3.8 27b Hanging in Pi.dev
LocalLLaMA
GotHereLateNameTaken18
🟠 redditQwen 3.8 27B OpenCode Config, With various reasoning effort low, med, xhigh, xhigh no preserve
LocalLLaMA
Altruistic_Heat_953160
🟠 redditAA is the reason for Qwen3.8 27B shipped with xhigh
LocalLLaMA
frontsideair20892
🟠 redditQwen 3.8 27B xhigh vs medium small comparison (+ others for fun)
LocalLLaMA
hiImMate11538
🟠 redditQwen 3.8 27b vs Deepseek Flash
LocalLLaMA
Best_Sail53257
🟠 redditQwen3.8 overthinks similarly (alot) on all reasoning eforts. Is it LM Studio bug?
LocalLLaMA
Jebbyk119
🟠 redditQwen 3.8 27b saved me $650+ in API costs this evening
LocalLLaMA
illgettheownerforyou115111
🟠 redditWe really need llama.cpp to support changing thinking amount
LocalLLaMA
TheWaffleKingg1833
🟠 redditLocal Qwen 3.8 27B vs GPT‑5.6 Terra vs Grok 4.6
artificial
Acceptable-Object39030
🟠 redditHow is qwen 3.8 slower than 3.6?
LocalLLaMA
Specialist-2193123
🟠 redditQwen 3.8 27B is the DeepSeek moment for local models. It matches frontier intelligence from just a few months ago and outperforms Google’s current frontier model. No big data center is needed as almost every serious local model expert can run it on their own hardware.
LocalLLaMA
InternationalGap369818999
🟠 redditAre you using --reasoning-preserve with llama.cpp and qwen3.8-27b?
LocalLLaMA
anderspitman716
🟠 redditPost-thinking sampler settings for vLLM
LocalLLaMA
TokenRingAI33
🟠 redditQwen 3.8 default temp (1.0) causes garbage output. Lowering to 0.1 fixes it, do I have something misconfigured?
LocalLLaMA
KeepyUpper017
🟠 redditMade this shooter game with Qwen 3.8 27B
LocalLLaMA
xdcfret111
🟠 redditA comparison of reasoning in two Qwen 3.8 releases
LocalLLaMA
k-r-a-u-s-f-a-d-r06
🟠 redditUnfortunately Qwen 3.8 27b is not good enough for complex coding
LocalLLaMA
myreala056
🟧 hnQwen3.8-27B make medium the default effort level instead of xhighxlayn133
🟠 redditQwen3.8-27B is disgustingly powerful
LocalLLaMA
bonobomaster4816
🟠 redditI ran Qwen3.8-27B on Mac M3 36G. 18 tok/s
LocalLLaMA
buryhuang28
🟠 redditWhy no "high" reasoning effort in Qwen 3.8 27b ?
LocalLLaMA
alanoo940
🟠 redditAm I doing something wrong? Qwen 3.8 27B seems useless for agentic coding
LocalLLaMA
BuahahaXD142226
🟠 redditupdated unsloth/Qwen3.8-27B-GGUF · Hugging Face
LocalLLaMA
jacek2023277111
🟠 redditReasoning can be broken for some altered qwen3.8 27bs
LocalLLaMA
fbms266
🟧 hnShow HN: 8-bit Qwen3.8-27B decodes 1.7x faster than BF16, slower at 16K contextgltanaka10
🟠 redditQwen 3.8 27B vs Gemini 3.7 Flash (High) for real coding: open-source 27B model did a much better job
LocalLLaMA
GravyPoo818
🟠 redditQwen 3.8 27b issues
LocalLLaMA
WyattTheSkid145
🟠 redditIntroducing Qwen3.8-27B Dynamic v3 Unsloth GGUFs
LocalLLaMA
danielhanchen1593240
🟠 redditQwen3.8-27B (Q5_K_XL) on Strix Halo at 31 t/s decode: DFlash2 + Vulkan, the optimal setup
LocalLLaMA
stereohype721
🟠 redditFluid Simulation Qwen3.8 27B IQ3_XXS
LocalLLaMA
Danmoreng3319
🟠 redditDFlash2 on 2x3090s - INT8 @ 140tps - 262k ctx
LocalLLaMA
luedtek2113
🟠 redditI pushed Qwen3.8-27B limits again... Dflash2 - 134 tps on a RTX 3090
LocalLLaMA
iamMess20479
🟠 redditQwen3.8-23B-Mini-Me: A Depth-Pruned Qwen3.8-27B (to ~22.7BB)
LocalLLaMA
peplo121415538
🟠 redditQwen 3.8 27B SlopCodeBench results
LocalLLaMA
corruptbytes4216
🟠 redditLarge Context w/ MTP, DFlash2, ngram-mod, Testing Qwen3.8-27B on 16GB VRAM
LocalLLaMA
BuffMcBigHuge1210
🟠 redditQwen3.8-27B took a serious hit to *knowledge* vs 3.6
LocalLLaMA
EmPips323225
🟠 redditI see some community discussions on Hugging Face about Qwen3.8 27B model being not very good, but it seems people are mostly have positive experiences here. So which is it?
LocalLLaMA
kr_tech077
🟠 redditI made Qwen 3.8 27B take the ACT to see if it’s ready for college.
LocalLLaMA
on_line1871814
🟠 redditQwen 3.8 27b *MEDIUM* is insane: 1/20th the thinking time of xhigh for almost the same quality output??
LocalLLaMA
9gxa05s8fa8sh215
🟠 redditQwen3.8-27B scored 29/30 on AIME 2026 with FP8 + xhigh reasoning β€” BF16 vs FP8 results
LocalLLaMA
No_Run88125149
🟠 redditGone back to 3.6 27B
LocalLLaMA
wsintra038
🟠 redditQwen3.8-27B at 262K context on a Strix Halo + RTX 3090 Ti: 9.5 -> 153 tok/s, and it beats a dual-3090 vLLM box on HumanEval
LocalLLaMA
TrifleHopeful54183943
🟠 redditI ran those benchmarks we all see on YouTube locally
LocalLLaMA
on_line18703
🟠 redditQwen 3.8 27b is strong even at Q3_xxs
LocalLLaMA
AltruisticList6000123116
🟠 redditIs qwen 3.8 27B actually qwen 3.6 27B who thinks (much) more?
LocalLLaMA
bajis12870011
🟠 redditQwen 3.8 vs 3.6 27b low reasoning loops way less now
LocalLLaMA
Lair982312
🟠 redditQwen3.8-27B on an RTX 5060 Ti 16GB: IQ4 vs Q8, 64K context, MTP, vision, and agent benchmarks
LocalLLaMA
Tema_Art_7777935
🟠 redditQwen3.8-27B Q6 is a beast at agentic coding
LocalLLaMA
Ok_Ninja7526543179
🟠 reddit3.8 reasoning for planning and instruct for applying the plan? Anyone tried it this way?
LocalLLaMA
soyalemujica610
🟠 redditQwen3.8-27B different thinking levels
LocalLLaMA
Tall_Abrocoma_353328663
🟠 redditQwen 3.8 Low and Medium are goated
LocalLLaMA
Eyelbee403125
🟠 redditWhich qwen 3.8 on rtx a2000?
LocalLLaMA
MrMrsPotts03
🟠 redditI'm really hoping we're in 2026's 2-month-gap between QwQ and Qwen3 right now
LocalLLaMA
ForsookComparison1020
🟠 redditI tried to do agenic coding with Qwen 3.8 27B 3bit quant on a macbook air m2 24gb. It took 63 hours, but amazingly, the flight simulator worked.
LocalLLaMA
HyperFoci15145
🟠 redditQwen 3.8 27B | First impressions
LocalLLaMA
CapsAdmin00
🟠 redditBro wtf, Qwen Lab cooked with Qwen 3.8 27B, it's so fucking good
LocalLLaMA
9r4n4y8857
🟠 redditWhy has my 3.6 35B become terrible now that I've started using 3.8 27B?
LocalLLaMA
N34257040
🟠 redditHelm chart for Qwen 3.8 for B70 users
LocalLLaMA
onebit10

Interpretation history

Decision trace