UkisAI, a small team behind creator Jovan Kis, open-sourced Swift-Qwen3.8-27B β a fine-tune of Qwen3.8-27B that attacks 'overthinking' by identifying reasoning-marker tokens ('wait', 'actually', 'let me reconsider') and penalizing them during training, then restoring accuracy via on-policy distillation. Creator-reported results are 58.3% fewer median thinking tokens on GPQA-Diamond at a 0.1-point accuracy cost and ~1.95x speedup; the model card itself discloses the hard-task exception (AIME 2026 falls 98.67%β94.00%), attributed to a penalized math-relevant training token with a fix promised. The release ships via Hugging Face with GGUF quants, vLLM/SGLang support, 262k context, and a free Nvidia-hosted API; the r/LocalLLaMA announcement drew heavy engagement where third-party paired benchmarks corroborate the token cuts but qualify the speed claim β decode throughput drops, so wall-clock gains come from fewer tokens plus faster prefill, not the 1.95x headline. The snippets cover the original 27B release; the Swift 1.5 family, the β63.4% headline, GSQ-RCO, and the 350k-downloads figure rest on the case's Reddit evidence rather than these web results β and adjacent hits (SAGE, ACL 2026's MUTO token-level marginal utility) show token-level overthinking penalization is an active, crowded training-research frontier.
2026-10-01T17:45:30Z
norenEnmotalen's matched-effort 69-question eval adding vanilla Unsloth Q4_K_XL and ThinkingCap alongside Swift1.5 closes the last methodological caveat (baseline-is-a-sibling-tune) on the third-party real-workload measurement β accuracy-per-token is now independently established across evaluators, hardware stacks and effort settings. With the periphery flat (same ~6 evaluators, one community, 6-12 pts per new thread vs 254-423 at peak) and the open questions collapsed to refinement level (AIME-fix confirmation, exact headline %s, 350k figure), the episode ends absorbed: proved out in corrected form (~33% token cuts, real wall-clock wins, mapped costs), not the creator's β63.4%/1.95x marketing frame.
2026-10-01T16:32:01Z
evidence attached: reddit.post.1wv1ico β Independent matched-effort head-to-head eval of Swift1.5 checkpoints against unsloth/ThinkingCap quants is material evidence for the Swift family's accuracy-per-token claim.
2026-10-01T00:58:05Z
The norenEnmotalen post adds a genuinely new evaluator running Swift 1.5 against a sibling Qwen fine-tune (Dirk-Qwen 3.8-27B) on personal domain workloads (M1 Max, 128k ctx) β the third-party real-workload paired measurement the case awaited, thickening the adoption ledger without answering any open question (AIME fix in 1.5, headline numbers, code cost, hard-task accuracy), and against another fine-tune rather than vanilla base. Engagement stays terminal (1.17 pts/h, 0.67 c/h, one community; the magnitude-valve spread reading remains inflated by the dead HN echo), so the quiet clock restarts on this datum and expire remains the natural next move unless the promised R9700 paired bench or an AIME-fix confirmation lands within days.
2026-10-01T00:29:45Z
evidence attached: reddit.post.1wuigui β Independent head-to-head domain evals of Swift-1.5 against Dirk-Qwen 3.8-27B is exactly the third-party real-workload measurement the Swift adoption case needs.
2026-09-29T19:17:24Z
The llm-bench.io M5 Max run extends the token-cut confirmation to Apple MLX (51k vs 77k generated tokens at equal quality) but is the same evaluator (DerTomsn) re-confirming the already-corroborated mechanism on new hardware β it thickens the ledger without moving any open question (AIME fix in 1.5, headline numbers, code cost, hard-task accuracy). The mild speedometer uptick (0.33β1.5 pts/h, coolingβsteady) tracks the new thin thread plus late votes, not renewed spread; magnitude-valve eligibility still overstates via the dead HN echo while in-community artifacts thin (4 pts/6 comments vs the 417/164 peak). Corroborated one-community token-efficiency win in terminal decay; quiet clock keeps running toward the promised R9700 paired bench or AIME-fix confirmation, and absent those within days, expire is the natural next move.
2026-09-29T17:41:59Z
evidence attached: reddit.post.1wte7n0 β Independent llm-bench.io run partially corroborates Swift 1.5's token-efficiency claim (51k vs 77k generated tokens at equal quality) on M5 Max.
2026-09-29T16:52:18Z
Tenth consecutive spike of identical shape: the 3.2x multiple is 3.17 pts/h on the KingGongzilla thread against a 1.0 peer baseline β another small-denominator artifact β plus comment churn treading settled ground on the two newest threads; no new facts, evaluators, artifacts, or spread beyond r/LocalLLaMA. Meaning is unchanged, and the speedometer has finally caught up to the standing judgment (0.33 pts/h vs 210 peak, momentum now reads cooling) β a corroborated one-community token-efficiency win with a filling-but-unresolved cost ledger in terminal decay, quiet clock running toward the promised R9700 paired Swift 1.5 bench or an AIME-fix confirmation.
2026-09-28T22:40:24Z
grounded: converges/high β Converges with what dev:project.llmreport's Agent Token Manifesto and his cost-tiered LiteLLM routing already argue β wasted thinking tokens are a real, attacka
2026-09-28T22:32:47Z
The KingGongzilla post adds the case's first task-level independent measurement β Swift 1.5 + HyperQwen on one RTX 3090, 37% less task-completion time at 100+ tps and 150k context β converting the corroborated token cut into wall-clock workload savings and extending the ecosystem to model+harness pairings, while the 360-run GLM prompt-discipline A/B gives weight-level training a prompt-level rival route, sharpening the case into a two-route token-discipline question (weights vs prompts) that maps directly onto Scott's harness-level thesis. Both are small, single-community posts amid ~2 pts/h terminal decay, so corroborated/low hold; the 'accelerating' momentum and magnitude-valve spread still overstate via small-denominator and dead-HN-echo artifacts, but the new independent measurement is material and restarts the quiet clock toward the promised R9700 paired bench or an AIME-fix confirmation.
2026-09-28T21:36:33Z
evidence attached: reddit.post.1wsnjzu β A 360-run A/B showing prompt-only 'thinking discipline' cutting GLM wasted thinking up to 70% is direct competitive context for whether trained efficiency models are the practical route.
2026-09-28T21:36:33Z
evidence attached: reddit.post.1wsqjku β Independent third-party task-level measurement of Swift 1.5 (+HyperQwen) showing 37% faster task completion is corroboration of real-workload adoption, not just download counts.
2026-09-28T06:46:05Z
Ninth consecutive velocity spike of identical shape: ~3.7 pts/h accruing on the newest thread (1wrv6bp) against a near-dead peer baseline β the 22x multiple is a small-denominator artifact β plus a few votes/comments on returnity re-treading settled points; no new facts, evaluators, artifacts, or spread beyond r/LocalLLaMA. Meaning is unchanged β a corroborated one-community token-efficiency win with a filling-but-unresolved cost ledger in terminal accretion decay β so corroborated/low hold; the 'steady' momentum and 80th-percentile peer reading overstate via age-cohort and dead-HN-echo artifacts, and the quiet clock keeps running toward the promised paired R9700 Swift 1.5 bench or an AIME-fix confirmation.
2026-09-28T03:43:44Z
Eighth sensor event is a small velocity bump on the newest thread (1wrv6bp: 3 pts/h off a tiny peer base β the 4.5x multiple is a small-denominator artifact) plus comment accretion re-treading known license/quant/code points, adding only anecdotal cost texture (a 7900xtx 'nothing meaningful', a code 'wash' on Go/Java); no new facts, artifacts, or cross-community spread. The episode remains a corroborated one-community win in terminal decay β case-wide ~4.8 pts/h vs 190 peak, 'steady' momentum only because the new bump masks old-thread decay β so corroborated/low hold, no material change, and the quiet clock keeps running toward the promised paired R9700 bench or an AIME-fix confirmation.
2026-09-27T21:50:59Z
Third independent evaluator adds a regime nuance β UkisAI's IQ4_XS beats Unsloth's Q4_K_S at low-thinking but loses at high-thinking β a thin, consistent datum that fills the quant-sensitivity question without changing the case's meaning; the seventh spike was again pure vote accretion (403β408, comments flat at 160), so corroborated/low hold and the quiet clock keeps running toward the promised paired R9700 bench or an AIME-fix confirmation.
2026-09-27T21:25:32Z
evidence attached: reddit.post.1wrv6bp β Independent community evaluation corroborating Swift 1.5's low-thinking quality edge over stock quants, plus a concrete practical success on a 3090.
2026-09-27T12:41:45Z
Sixth consecutive velocity spike of identical shape: ~16 more votes on the sleight42 Swift 1.5 thread (~379β~395, 155β158 comments) over ~7h at ~0.67 comments/h, with no new evidence, artifacts, or substantive comments β the periphery has been static since the abliterated quant, and the 96.9th-percentile spread reading is an age-cohort artifact (a 5 pts/h thread at 68h looks fast only against decayed peers) resting on a dead 1/1 HN echo. Meaning unchanged β corroborated token savings with a filling-but-unresolved cost ledger in pure accretion tail β so corroborated/low hold; the quiet clock should run pending the promised paired R9700 Swift 1.5 bench or an AIME-fix confirmation.
2026-09-27T04:35:52Z
Fifth consecutive velocity spike of identical shape: votes accruing on the sleight42 Swift 1.5 thread (~348β379) while the carrier ticks DOWN (312β305) and discussion flatlines (~0.3 comments/h) β pure accretion on already-counted evidence, no new facts, no new artifacts, periphery static. The magnitude-valve spread reading still overstates via the dead 1/1 HN echo, so heat holds at low; the case's meaning is unchanged β corroborated token savings with a filling-but-unresolved cost ledger, one-community episode in its terminal cooling tail.
2026-09-26T21:41:58Z
Fourth consecutive spike of the same shape: votes accumulating on the sleight42 Swift 1.5 thread (~238β~348 in ~6h) plus carrier creep (307β312), with comments re-treading known license/quant/derivative points and no new facts; the periphery itself stopped expanding this window (no new artifacts since the abliterated quant), so the magnitude valve's spread reading still rests on a dead HN echo. Meaning unchanged β corroborated token savings with a filling-but-unresolved cost ledger, one-community episode in its steady cooling tail β so corroborated/low hold.
2026-09-26T15:46:37Z
The velocity spike is point accumulation on the sleight42 Swift 1.5 thread (212β238 in ~90 min), the third consecutive spike of this shape β engagement thickening inside r/LocalLLaMA, not new information; the magnitude valve's 'second platform' remains the score-1 HN echo. The case's meaning is unchanged (corroborated token savings with a filling cost ledger, one-community episode in its steady tail), so corroborated/low hold: the alert already delivered and the open questions (AIME fix in 1.5, code-cost reproduction, quant-sensitivity, 350k figure) resolve on day timescales, not hours.
2026-09-26T12:41:51Z
The anticipated R9700-class reproduction partially arrived as a first-hand report (dual R9700, FP8 + DFlash: ~5100 t/s prefill, 130+ t/s decode, positive code-gen), reinforcing the practical-usability leg but without a paired base-model baseline and with speculative decoding confounding throughput β so the decode-drop, code-cost, and AIME-fix questions all stand. The case's meaning is unchanged (corroborated token savings with a filling cost ledger), now clearly in the cooling tail of a one-community episode; heat steps to low because nothing in the open-questions list resolves on hour timescales.
2026-09-26T08:28:22Z
Since the code-cost datum, only the known periphery thickened: a first third-party derivative artifact (abliterated NInfer nvfp4 quant) and more appreciation threads within r/LocalLLaMA, plus a second quant-confounded anecdote that xhigh savings shrink on nvfp4 β ecosystem-forming and fragility-reinforcing, but no new fact that changes the assessment. The returnity thread's velocity spike is point accumulation on already-counted evidence; still one community, cooling (10 pts/h vs 173 peak), so corroborated/medium hold and this look is not material.
2026-09-26T08:22:35Z
evidence attached: reddit.post.1wqkb3v β Swift 1.5 is a direct follow-up release on the open corroborated Swift-family case, with community quant derivatives already appearing β momentum evidence a re-judge would need.
2026-09-26T06:46:47Z
The cost ledger gained its first quantified independent datum: a community member reports a 16-point drop on a selective code benchmark for Swift 1.5 Flash vs base Flash (nvfp4, single source, quant confound possible), converting the earlier 'python proficiency' anecdote into a testable counter-signal β the case's meaning firms up as 'real, corroborated token savings with a filling cost ledger', not a free lunch. Otherwise flat: engagement barely moved and is cooling (7.2 pts/h vs 135 peak), still effectively one community, so heat steps down from the magnitude-valve high β the alert already delivered, periphery expands only within r/LocalLLaMA, and nothing here warrants within-hours rechecks.
2026-09-25T19:50:11Z
magnitude valve eligible (multi-platform, top-decile engagement) and never alerted; deterministic escalation to deliver
2026-09-25T19:24:46Z
evidence attached: reddit.post.1wq56pf β First-person community adoption report of the new Swift1.5-Qwen3.8-Flash-Next variant, corroborating the Swift family's sustained-adoption and overthinking-reduction claims.
2026-09-25T00:12:12Z
First independent corroboration landed: a third-party paired benchmark (7900 XTX) confirms the core mechanism β Swift cuts ~33% of tokens vs the base Qwen3.8-27B quant across four scenarios, beating rival low-thinking quant ThinkingCap's β26% β while qualifying the speed claim: Swift's decode throughput drops, so wall-clock gains come from fewer tokens plus much faster prefill, not the 1.95x headline. The case moves announcement-class β corroborated; headline β63%/1.95x and the 350k downloads figure remain creator-reported.
2026-09-24T23:33:56Z
evidence attached: reddit.post.1wpg32w β Independent 7900 XTX benchmark corroborates Swift's ~33% token cut with quality held, while qualifying the speed claim (decode β33%, much faster prefill).
2026-09-24T23:06:29Z
Same announcement, same claims β what changed is the shape of the wave and the first independent flickers: Reddit carries the episode at top-decile rate (92.5 peer percentile, steady off a ~56 pts/h peak) while the company's own HN cross-post died at score 1, so this is one strong platform plus a weak echo, not broad multi-platform dominance; comments surfaced the first faint independent usage anecdote (two weeks of 27B use, homelab sysadmin) and a concrete adoption friction (license restriction layered on Qwen pushed at least one HN user elsewhere). Zero independent verification of the β63%/1.95x benchmarks or the 350k figure β announcement-class with strong traction, and engagement alone doesn't buy corroborated.
2026-09-24T20:37:40Z
evidence attached: hn.story.49833102 β Same UkisAI Swift-family release announcement surfacing on HN; company-posted so adoption context, not independent corroboration.
2026-09-24T18:47:28Z
grounded: converges/high β Dated-receipts convergence: UkisAI bakes into Qwen3.8-27B weights the thesis Scott argues at the harness/orchestration level β that maximal per-step reasoning i
2026-09-24T18:39:19Z
case created β First-party release announcement from the claim owner with strong traction and a resolvable adoption-plus-efficiency claim, distinct from the narrower ternary Bonsai 2 quant case already open.