2026-10-11 16:38 UTC

Redditor MushroomMan234 reports that UkisAI's Swift 1.5 โ€” a reasoning-efficient fine-tune of Qwen3.8-Flash-Next โ€” running on sf-stav's veloGB10 GB10-only engine sustains ~110 tok/s decode on two DGX Sparks (vs ~52 for base NVFP4 on vLLM) and beats base Flash-Next at medium effort on coding-agent pass rate (92% vs 50%), making a fine-tune-plus-single-model-engine stack a demonstrated local coding-agent path on GB10 hardware if others replicate it.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumlocal-inference dgx-spark agentic-coding fine-tuningUkisAIsf-stav

What is this?

The case rests on a Reddit report that Swift 1.5, a reasoning-efficient fine-tune of NVIDIA's Qwen3.8-Flash-Next by a lab called UkisAI, sustains ~110 tok/s and much higher coding-agent pass rates on two DGX Sparks when served by sf-stav's veloGB10. The stack is well corroborated: veloGB10 is a real, actively developed from-scratch Rust+CUDA inference engine built specifically for NVIDIA's GB10 (DGX Spark), serving Flash-Next from NVFP4 and EXL3 packs, and the NVIDIA DGX Spark forums show a dense measurement culture around Flash-Next with single-stream baselines roughly 43โ€“64 tok/s on one Spark (veloGB10's own v0.7.0 release claims 186 tok/s TP=2 for base Flash-Next in EXL3). However, no supplied snippet mentions UkisAI, Swift 1.5, or the specific ~110 tok/s / 92%-vs-50% figures โ€” the closest fragment is a forum comment crediting an unnamed 'INT4 custom model' with ~92 on tool-eval-bench at medium effort and much faster speed, which rhymes with but does not confirm the fine-tune claim, so the headline numbers remain single-source and unreplicated.

Why it matters to Scott

The stack converges with his hardware-aware-local-inference thesis โ€” veloGB10 is a from-scratch single-model engine treating GB10 placement and NVFP4/EXL3 precision as explicit runtime policy โ€” and if the 92-vs-50 finetune-over-base result replicates it would be the first community-finetune win against the radar's returnity baseline (five Qwen3.6-35B finetunes, none beat base on coding), bearing directly on whether his synthetic-finetuning data-factory work has agentic payoff. Until replication it stays a single-source Reddit claim, so it extends the DGX Spark / single-model-engine lineage rather than settling anything.
dev:concept.hardware-aware-local-inferencedev:concept.synthetic-finetuning-datasetradar:qwen36-35b-finetunes-vs-baseradar:qwen38-flashnext-custom-engineradar:dual-dgx-spark-deepseek-flash-v4radar:nvidia-spark-line-repricingradar:concept.dgx-sparkradar:concept.local-inference
queries asked of Scott's wikis
  • dgx spark gb10 local agent hardware
  • local coding agent model throughput usability threshold
  • fine-tune vs base model agentic coding performance
  • reasoning effort latency cost agent loops
  • nvfp4 exl3 quantization weight format tradeoffs
  • single-model custom inference engine vs vllm

Measured heat

now 0 pts/hpeak 20 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 145h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-05 15:04โญ origin directly observedSwift 1.5 on veloGB10, ~110 tok/s on 2ร— DGX Spark: xhigh beats base Flash-Next at medium on vLLM
MushroomMan234 on r/LocalLLaMA
โ€”
10-05 15:04amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wyaw0o
MushroomMan234
peak 9 ยท 19 comments ยท 100% of case engagement
10-05 15:20our radar first saw it ยท +0.3hdiscovery anchor: reddit.post.1wyaw0oโ€”
pace: p57 vs 1247 stories at the 96h mark (now 145h old) โ€” ahead of anthropic-fourth-cyber-incident-review-miss (1.0x), behind backburner-iphone-offload (1.0x)

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญSwift 1.5 on veloGB10, ~110 tok/s on 2ร— DGX Spark: xhigh beats base Flash-Next at medium on vLLM
LocalLLaMA
MushroomMan234919

Interpretation history

Decision trace