LocalLLaMA user returnity's comparison of five Qwen3.6-35B-A3B community finetunes finds none beats the base model on coding evaluations (only Occamy-1.0 competitive), and wider replication โ or a finetune that clearly wins โ resolves whether community finetunes add real value over base for small-MoE local workflows.
state: resolvedheat: lowuncertainty: mediumnovelscott: mediumlocal-models model-evaluation local-inferencereturnity
What is this?
Qwen3.6-35B-A3B is Alibaba's open-weight sparse MoE (35B total parameters, ~3B active per token, Apache 2.0), released April 16, 2026 and positioned as a workstation-class agentic coding model โ vendor-reported 73.4 SWE-bench Verified and 51.5 Terminal-Bench 2.0, runnable from ~20-22GB Q4 quants. An active community-finetune ecosystem has formed around it (e.g. the Occamy-1.0 finetune ships its own benchmark table claiming large gains over base), r/LocalLLaMA routinely swaps comparative evals of these checkpoints, and commenters are openly awaiting a Qwen 3.8 refresh. The search results establish that model and ecosystem but did not surface returnity's specific post; per the case, it tests five 35B-A3B finetunes against base on coding tasks and finds none beats base, with Occamy-1.0 the only competitive one. One early writeup called the 3.6 release 'plausible but uncorroborated' against verified Qwen3.5 artifacts, though multiple independent sources and Hugging Face repos now treat 3.6 as shipped.
Why it matters to Scott
A credible community null-result โ none of five finetunes beats base Qwen3.6-35B-A3B on coding โ bears on the premise of Scott's own finetuning-data work (his reddit/Salesforce data factory and synthetic-dataset concept only pay off if a finetune beats base+prompting) and gives a concrete steer for which checkpoints belong in the gamepc zoo: serve base for agentic coding until a finetune clearly wins. It also sharpens open radar questions โ Surge AI's claimed +5.8pp SWE-Bench Pro from post-training and the 16GB Qwen3.6 LoRA training workflow โ though the result is a single community eval pending replication. Scott's canon holds no prior position on finetune-vs-base value, so this is new evidence rather than a convergence or challenge.
dev:concept.synthetic-finetuning-datasetdev:project.redditdev:project.gamepcradar:concept.fine-tuningradar:gguf-lora-16gb-moe-trainingradar:surge-office-training-coding-transfer
queries asked of Scott's wikis
- community finetune vs base model value for coding agents
- sparse MoE active-parameter economics local workstation inference
- local model selection beyond vendor benchmark tables
- running local open models inside coding agent harnesses
- when is LoRA finetuning worth it vs base model plus prompting
- small model ceiling for agentic tool use and long context
Measured heat
no measured readings yet โ the hourly heat pass fills this in
How the heat travelled
Evidence (1) โ โญ canonical anchor
Interpretation history
2026-10-02T17:01:04Z
The wider replication the hypothesis called for arrived: OsmanthusBloom's independent tool-calling benchmark shows Tiel and Ornith clearly beating base Qwen, while returnity's Aider Polyglot eval shows none do โ so the verdict flips from 'contested null result' to a settled, community-level answer: finetune value is harness/task-contingent, real on tool calling, absent on Aider-style agentic coding. The episode's question is answered as well as this ecosystem will answer it; the thread is inert (0 pts/h, single platform, ~91h old) and the expected Qwen 3.8 refresh would be a new generation, not a continuation.
2026-09-30T09:59:04Z
The null-result is now benchmark-contingent rather than clean: the Tiel quant author attributes its Q4 edge to dynamic imatrix calibration (absent at Q8), a practitioner reports base underperforming Ornith/Tiel on real repos, and commenters dispute Aider Polyglot's agentic validity โ so the case reads as a single eval with identified confounds pending replication, still a useful steer but weaker.
2026-09-29T00:57:28Z
grounded: novel/medium โ A credible community null-result โ none of five finetunes beats base Qwen3.6-35B-A3B on coding โ bears on the premise of Scott's own finetuning-data work (his r
2026-09-29T00:49:43Z
case created โ A substantive comparative evaluation with real comment engagement and a decision-relevant, refutable claim about which checkpoints local builders should actually deploy.
Decision trace
- 10-03 03:01resolveThe wider replication the hypothesis called for arrived: OsmanthusBloom's independent tool-calling benchmark shows Tiel and Ornith clearly beating base Qwen, while returnity's Aider Polyglot
- 10-03 02:59jev_reprice_gatechanges_anything noul=0.60 would_skip=False
- 10-03 02:59review_screenNew comment by OsmanthusBloom (replacing peculiar-ragdoll's slot) cites their own separate tool-calling benchmark where Tiel and Ornith clearly beat base Qwen, implying the finetunes are weak onl
- 10-03 02:58review_screenjev screen borderline (noul=0.72) โ luna review
- 09-30 19:59repriceThe null-result is now benchmark-contingent rather than clean: the Tiel quant author attributes its Q4 edge to dynamic imatrix calibration (absent at Q8), a practitioner reports base underperforming O
- 09-29 23:21sensor_dirtycomment_update
- 09-29 16:21sensor_dirtycomment_update
- 09-29 12:22sensor_dirtyvelocity_spike
- 09-29 10:57groundA credible community null-result โ none of five finetunes beats base Qwen3.6-35B-A3B on coding โ bears on the premise of Scott's own finetuning-data work (his reddit/Salesforce data factory and s
- 09-29 10:49createA substantive comparative evaluation with real comment engagement and a decision-relevant, refutable claim about which checkpoints local builders should actually deploy.