Reddit builder Few-Rough-2215 reports that Unsloth's LoRA merge path applies roughly double the nominal alpha/r effective scale โ silently discarding most of the tuning on nominal merges โ and independent reproduction or an Unsloth fix or acknowledgment decides whether local fine-tuners' nominal merges are systematically mis-scaled.
state: seedheat: lowuncertainty: mediumconvergesscott: highlora-fine-tuning unsloth local-model-toolingFew-Rough-2215
What is this?
Reddit builder Few-Rough-2215 โ who fine-tuned MedGemma 4B (LoRA) and 27B (QLoRA) for oncology on a single NVIDIA DGX Spark โ reports a suspected scaling bug in Unsloth's LoRA merge path: merged models appear to apply roughly double the nominal alpha/r effective scale, so a nominal (alpha=r) merge could silently mis-carry most of the tuning. The claim rests on the builder's own paired adapter-vs-merge stats and stays unconfirmed until independently reproduced, fixed, or acknowledged by Unsloth. The supplied snippets do not surface the post itself or any Unsloth response, but the terrain is credible: Unsloth's own docs pin alpha/r=1 as baseline (alpha=2r as an 'aggressive' heuristic), and merge-path defects are already a reported failure class there โ GitHub issue #5410 shows `save_pretrained_merged` producing garbage output while the same adapters loaded separately work fine, and an r/unsloth thread reports merged-vs-dynamic behavior discrepancies. So this is a checkable tooling-bug claim against the dominant local fine-tuning library: plausible given prior merge-path complaints, but the specific 2x mechanism is unresolved.
Why it matters to Scott
This lands on Scott's own toolchain, not just his themes: the gamepc wiki carries Unsloth/PEFT LoRA pipeline notes and a merge-vs-adapter parity eval practice, and his domain fine-tune projects (the reddit data factory, synthetic fine-tuning datasets) sit squarely in the affected population โ if the 2x claim reproduces, merged exports need re-audit and adapter-at-inference becomes the safer default. It independently arrives where his canon already argues: the builder's paired adapter-vs-merge stats are exactly the deterministic-gate parity check Scott practices, and the finding hands that practice a named first-suspect (alpha/r scale handling) to check before deeper numerics whenever merge-vs-adapter divergence shows up.
dev:project.gamepcdev:project.redditdev:concept.synthetic-finetuning-datasetdev:concept.round-trip-fidelity-scoringradar:person.unslothradar:concept.loraradar:concept.fine-tuningradar:unsloth-gpt-oss-template-bugradar:concept.reproducibility
queries asked of Scott's wikis
- Unsloth or PEFT LoRA fine-tuning pipeline dev notes
- LoRA merge vs adapter-at-inference parity eval practice
- silent numerics or scaling bug in ML toolchain reproduction protocol
- DGX Spark local training workstation projects
- fine-tune vs RAG for domain knowledge tradeoff position
- local open model customization strategy notes
Measured heat
now 0 pts/hpeak 2 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 128h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p9 vs 1247 stories at the 96h mark (now 128h old) โ behind 3jsbench-llm-3d-generation-benchmark (0.5x)
Evidence (1) โ โญ canonical anchor
Interpretation history
2026-10-06T09:46:34Z
grounded: converges/high โ This lands on Scott's own toolchain, not just his themes: the gamepc wiki carries Unsloth/PEFT LoRA pipeline notes and a merge-vs-adapter parity eval practice,
2026-10-06T09:38:33Z
case created โ Concrete checkable tooling-bug claim backed by paired stats in a dominant local fine-tuning library, a different Unsloth episode from the open GPT-OSS template-bug case.
Decision trace
- 10-06 20:46groundThis lands on Scott's own toolchain, not just his themes: the gamepc wiki carries Unsloth/PEFT LoRA pipeline notes and a merge-vs-adapter parity eval practice, and his domain fine-tune projects (
- 10-06 20:38createConcrete checkable tooling-bug claim backed by paired stats in a dominant local fine-tuning library, a different Unsloth episode from the open GPT-OSS template-bug case.