2026-10-11 17:13 UTC

UkisAI claims its Swift family of Qwen-derived reasoning models cuts pathological overthinking tokens by ~63% at ~1.95x speed with accuracy restored via GSPO/OPD training, and its 350k+ downloads in 13 days mark sustained adoption as a practical accuracy-per-token option for local efficient reasoning.

state: resolvedheat: lowuncertainty: lowconvergesscott: highlocal-models efficient-reasoning model-releases inference-economicsUkisAIJovan
Surfaced 2026-09-25T19:59:25Z β€” UkisAI Swift Series / 27B, Flash Next and Bonsai 2 + GSQ-RCO / -63.4% thinking, x1.95 speed with xhigh accuracy β€” Corroboration broadened from a single benchmark to an ecosystem signal: a second independent community voice (returnity, with its own earlier Swift/ThinkingCap/base comparison) now adopts the family per-variant, reporting Swift1.5-Flash-Next as a clear win over base Flash, while a new quality-cost question surfaced (users asking why Swift seems to impact python proficiency) and the promised AIME-regression fix remains unanswered. Spread is still effectively one community (HN echo dead), so corroborated holds rather than accelerating; heat stays medium because the carrier is past peak (4.7 pts/h vs 65.8 peak) yet the periphery keeps expanding β€” new bench threads, per-variant adoption, planned R9700 reproduction, Unsloth GGUF demand β€” and the magnitude-valve spread reading overstates cross-platform presence (second platform is a score-1 echo).

What is this?

UkisAI, a small team behind creator Jovan Kis, open-sourced Swift-Qwen3.8-27B β€” a fine-tune of Qwen3.8-27B that attacks 'overthinking' by identifying reasoning-marker tokens ('wait', 'actually', 'let me reconsider') and penalizing them during training, then restoring accuracy via on-policy distillation. Creator-reported results are 58.3% fewer median thinking tokens on GPQA-Diamond at a 0.1-point accuracy cost and ~1.95x speedup; the model card itself discloses the hard-task exception (AIME 2026 falls 98.67%β†’94.00%), attributed to a penalized math-relevant training token with a fix promised. The release ships via Hugging Face with GGUF quants, vLLM/SGLang support, 262k context, and a free Nvidia-hosted API; the r/LocalLLaMA announcement drew heavy engagement where third-party paired benchmarks corroborate the token cuts but qualify the speed claim β€” decode throughput drops, so wall-clock gains come from fewer tokens plus faster prefill, not the 1.95x headline. The snippets cover the original 27B release; the Swift 1.5 family, the βˆ’63.4% headline, GSQ-RCO, and the 350k-downloads figure rest on the case's Reddit evidence rather than these web results β€” and adjacent hits (SAGE, ACL 2026's MUTO token-level marginal utility) show token-level overthinking penalization is an active, crowded training-research frontier.

Why it matters to Scott

Converges with what dev:project.llmreport's Agent Token Manifesto and his cost-tiered LiteLLM routing already argue β€” wasted thinking tokens are a real, attackable cost β€” but the new material converts agreement into a fork his harness-level thesis must now answer: a 360-run GLM A/B cuts wasted thinking ~70% with prompt rules alone (the control experiment for weight-level retraining), while KingGongzilla's 37% task-time reduction is the first wall-clock receipt that token cuts survive real workloads. It is directly actionable in his own stack β€” a paired Swift 1.5 vs base Qwen3.8-27B run in the gamepc/Ollama zoo would decide whether the cheap tier of his routing upgrades to an efficient-reasoning model or a disciplined brief β€” and the weights-vs-prompts comparison is dated-receipt material for the Manifesto's token-efficiency territory.
dev:project.llmreportdev:technology.litellmdev:concept.cost-tiered-llm-routingdev:technology.ollamadev:project.gamepcradar:concept.token-efficiencyradar:concept.inference-economicsradar:concept.qwen38radar:qwen-hedging-token-logit-biasradar:qwen38-27b-16gb-quant-benchmarkradar:karpathy-claude-md-rules
queries asked of Scott's wikis
  • mature token law wasted thinking tokens
  • harness-level vs weight-level token discipline reasoning efficiency
  • overthinking loops coding agents context budget waste
  • local inference economics accuracy per token
  • on-policy distillation GSPO Qwen fine-tune
  • prompt rules to cut reasoning length

Measured heat

no measured readings yet β€” the hourly heat pass fills this in

How the heat travelled

09-24 16:32⭐ origin directly observedUkisAI Swift Series / 27B, Flash Next and Bonsai 2 + GSQ-RCO / -63.4% thinking, x1.95 speed with xhigh accuracy
Secure_Recording_472 on r/LocalLLaMA
β€”
09-24 16:33first on hacker news Β· published Β· +0.0hUkisAI Swift Series / 27B, Flash Next and Bonsai 2 /-63.4% thinking, x1.95 speed
kisjovan
β€”
09-24 22:49first on r/LocalLLaMA Β· published Β· +6.3h7900 XTX β€” two "low-thinking" Qwen 3.8 27B quants (Swift + ThinkingCap) vs the regular quant
DerTomsn
β€”
09-24 16:32amplified on r/LocalLLaMAreddit.post.1wp6gal
Secure_Recording_472
peak 317 Β· 250 comments Β· 32% of case engagement
09-24 16:33amplified on hacker newshn.story.49833102
kisjovan
peak 1 Β· 1 comments Β· 0% of case engagement
09-24 22:49amplified on r/LocalLLaMAreddit.post.1wpg32w
DerTomsn
peak 6 Β· 15 comments Β· 1% of case engagement
09-25 19:16amplified on r/LocalLLaMAreddit.post.1wq56pf
returnity
peak 257 Β· 108 comments Β· 20% of case engagement
09-26 07:28amplified on r/LocalLLaMA πŸ‘‘reddit.post.1wqkb3v
sleight42
peak 426 Β· 166 comments Β· 33% of case engagement
09-27 20:46amplified on r/LocalLLaMAreddit.post.1wrv6bp
MomentJolly3535
peak 56 Β· 35 comments Β· 5% of case engagement
5 more amplifiers in ainews.case_chain
09-24 18:20our radar first saw it Β· +1.8hdiscovery anchor: reddit.post.1wp6galβ€”
09-25 19:50reached heat=high Β· +27.3h Β· via ledgerβ€”β€”

Evidence (11) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐UkisAI Swift Series / 27B, Flash Next and Bonsai 2 + GSQ-RCO / -63.4% thinking, x1.95 speed with xhigh accuracy
LocalLLaMA
Secure_Recording_472315252
🟧 hnUkisAI Swift Series / 27B, Flash Next and Bonsai 2 /-63.4% thinking, x1.95 speedkisjovan11
🟠 reddit7900 XTX β€” two "low-thinking" Qwen 3.8 27B quants (Swift + ThinkingCap) vs the regular quant
LocalLLaMA
DerTomsn515
🟠 redditSwift1.5-Qwen3.8-Flash-Next is phenomenal vs. base 3.8-Flash!
LocalLLaMA
returnity254108
🟠 redditSwift 1.5 27b: Swift Qwen just got faster
LocalLLaMA
sleight42423166
🟠 redditSwift 1.5 Qwen3.8 27b (A must-have for low thinking!)
LocalLLaMA
MomentJolly35355835
🟠 redditSwift 1.5 + HyperQwen = 37% less task completion time at 100+ tps w/ 150k context on RTX 3090
LocalLLaMA
KingGongzilla3818
🟠 reddit9 prompt rules cut my coding agent's wasted thinking up to 70% (GLM 5.3 & GLM 5.3 Flash)
LocalLLaMA
PilgrimofHaqq23320
🟠 redditSwift-1.5-Qwen3.8-27b-oQ8e-mtp on Apple M5 Max β€” 34.8 tok/s β€” llm-bench.io
LocalLLaMA
DerTomsn87
🟠 redditpeculiar-ragdoll's Dirk-Qwen 3.8-27B vs. UkisAI Swift-1.5 Qwen3.8-27B
LocalLLaMA
norenEnmotalen1320
🟠 redditUnsloth, Swift1.5, Peculiar-Ragdoll, ThinkingCap - Qwen3.8-27B
LocalLLaMA
norenEnmotalen2122

Interpretation history

Decision trace