2026-10-11 17:10 UTC

fine-tuning

band: coolmomentum: stable score: 0.125
temperature history

Episodes (4)

Independent evaluation will determine whether fine-tuning a 450M-parameter vision-language model on 50,000 browser screenshots raises held-out browser-interface understanding from 1% to roughly 44% while preserving a substantial efficiency advantage.
expiredknownscott: medium
Independent reproduction will determine whether supervised fine-tuning with Qwen3โ€™s default chat template can suppress thinking mode while leaving training loss and basic evaluations apparently normal.
expirednovelscott: low
tamewild claims its released PCSS recipe fine-tunes Qwen3-4B Base on 100 zebra puzzles in roughly 6.5 GPU minutes to reach 85.26% on the full MATH benchmark, a reported 31.16-point gain that could make substantial small-model reasoning improvements unusually cheap.
seedconvergesscott: low
Redditor MushroomMan234 reports that UkisAI's Swift 1.5 โ€” a reasoning-efficient fine-tune of Qwen3.8-Flash-Next โ€” running on sf-stav's veloGB10 GB10-only engine sustains ~110 tok/s decode on two DGX Sparks (vs ~52 for base NVFP4 on vLLM) and beats base Flash-Next at medium effort on coding-agent pass rate (92% vs 50%), making a fine-tune-plus-single-model-engine stack a demonstrated local coding-agent path on GB10 hardware if others replicate it.
seedconvergesscott: medium

Trajectory notes