2026-10-11 16:37 UTC

Reddit builder Few-Rough-2215 reports that Unsloth's LoRA merge path applies roughly double the nominal alpha/r effective scale โ€” silently discarding most of the tuning on nominal merges โ€” and independent reproduction or an Unsloth fix or acknowledgment decides whether local fine-tuners' nominal merges are systematically mis-scaled.

state: seedheat: lowuncertainty: mediumconvergesscott: highlora-fine-tuning unsloth local-model-toolingFew-Rough-2215

What is this?

Reddit builder Few-Rough-2215 โ€” who fine-tuned MedGemma 4B (LoRA) and 27B (QLoRA) for oncology on a single NVIDIA DGX Spark โ€” reports a suspected scaling bug in Unsloth's LoRA merge path: merged models appear to apply roughly double the nominal alpha/r effective scale, so a nominal (alpha=r) merge could silently mis-carry most of the tuning. The claim rests on the builder's own paired adapter-vs-merge stats and stays unconfirmed until independently reproduced, fixed, or acknowledged by Unsloth. The supplied snippets do not surface the post itself or any Unsloth response, but the terrain is credible: Unsloth's own docs pin alpha/r=1 as baseline (alpha=2r as an 'aggressive' heuristic), and merge-path defects are already a reported failure class there โ€” GitHub issue #5410 shows `save_pretrained_merged` producing garbage output while the same adapters loaded separately work fine, and an r/unsloth thread reports merged-vs-dynamic behavior discrepancies. So this is a checkable tooling-bug claim against the dominant local fine-tuning library: plausible given prior merge-path complaints, but the specific 2x mechanism is unresolved.

Why it matters to Scott

This lands on Scott's own toolchain, not just his themes: the gamepc wiki carries Unsloth/PEFT LoRA pipeline notes and a merge-vs-adapter parity eval practice, and his domain fine-tune projects (the reddit data factory, synthetic fine-tuning datasets) sit squarely in the affected population โ€” if the 2x claim reproduces, merged exports need re-audit and adapter-at-inference becomes the safer default. It independently arrives where his canon already argues: the builder's paired adapter-vs-merge stats are exactly the deterministic-gate parity check Scott practices, and the finding hands that practice a named first-suspect (alpha/r scale handling) to check before deeper numerics whenever merge-vs-adapter divergence shows up.
dev:project.gamepcdev:project.redditdev:concept.synthetic-finetuning-datasetdev:concept.round-trip-fidelity-scoringradar:person.unslothradar:concept.loraradar:concept.fine-tuningradar:unsloth-gpt-oss-template-bugradar:concept.reproducibility
queries asked of Scott's wikis
  • Unsloth or PEFT LoRA fine-tuning pipeline dev notes
  • LoRA merge vs adapter-at-inference parity eval practice
  • silent numerics or scaling bug in ML toolchain reproduction protocol
  • DGX Spark local training workstation projects
  • fine-tune vs RAG for domain knowledge tradeoff position
  • local open model customization strategy notes

Measured heat

now 0 pts/hpeak 2 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 128h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-06 07:55โญ origin directly observedFine-tuned MedGemma 4B (LoRA) and 27B (QLoRA) for oncology on one DGX Spark. Also: a possible LoRA scale discrepancy under Unsloth, looking for independent reproduction
Few-Rough-2215 on r/LocalLLaMA
โ€”
10-06 07:55amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wywscq
Few-Rough-2215
peak 6 ยท 0 comments ยท 100% of case engagement
10-06 08:20our radar first saw it ยท +0.4hdiscovery anchor: reddit.post.1wywscqโ€”
pace: p9 vs 1247 stories at the 96h mark (now 128h old) โ€” behind 3jsbench-llm-3d-generation-benchmark (0.5x)

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญFine-tuned MedGemma 4B (LoRA) and 27B (QLoRA) for oncology on one DGX Spark. Also: a possible LoRA scale discrepancy under Unsloth, looking for independent reproduction
LocalLLaMA
Few-Rough-221560

Interpretation history

Decision trace