2026-10-11 17:21 UTC

LocalLLaMA user returnity's comparison of five Qwen3.6-35B-A3B community finetunes finds none beats the base model on coding evaluations (only Occamy-1.0 competitive), and wider replication โ€” or a finetune that clearly wins โ€” resolves whether community finetunes add real value over base for small-MoE local workflows.

state: resolvedheat: lowuncertainty: mediumnovelscott: mediumlocal-models model-evaluation local-inferencereturnity

What is this?

Qwen3.6-35B-A3B is Alibaba's open-weight sparse MoE (35B total parameters, ~3B active per token, Apache 2.0), released April 16, 2026 and positioned as a workstation-class agentic coding model โ€” vendor-reported 73.4 SWE-bench Verified and 51.5 Terminal-Bench 2.0, runnable from ~20-22GB Q4 quants. An active community-finetune ecosystem has formed around it (e.g. the Occamy-1.0 finetune ships its own benchmark table claiming large gains over base), r/LocalLLaMA routinely swaps comparative evals of these checkpoints, and commenters are openly awaiting a Qwen 3.8 refresh. The search results establish that model and ecosystem but did not surface returnity's specific post; per the case, it tests five 35B-A3B finetunes against base on coding tasks and finds none beats base, with Occamy-1.0 the only competitive one. One early writeup called the 3.6 release 'plausible but uncorroborated' against verified Qwen3.5 artifacts, though multiple independent sources and Hugging Face repos now treat 3.6 as shipped.

Why it matters to Scott

A credible community null-result โ€” none of five finetunes beats base Qwen3.6-35B-A3B on coding โ€” bears on the premise of Scott's own finetuning-data work (his reddit/Salesforce data factory and synthetic-dataset concept only pay off if a finetune beats base+prompting) and gives a concrete steer for which checkpoints belong in the gamepc zoo: serve base for agentic coding until a finetune clearly wins. It also sharpens open radar questions โ€” Surge AI's claimed +5.8pp SWE-Bench Pro from post-training and the 16GB Qwen3.6 LoRA training workflow โ€” though the result is a single community eval pending replication. Scott's canon holds no prior position on finetune-vs-base value, so this is new evidence rather than a convergence or challenge.
dev:concept.synthetic-finetuning-datasetdev:project.redditdev:project.gamepcradar:concept.fine-tuningradar:gguf-lora-16gb-moe-trainingradar:surge-office-training-coding-transfer
queries asked of Scott's wikis
  • community finetune vs base model value for coding agents
  • sparse MoE active-parameter economics local workstation inference
  • local model selection beyond vendor benchmark tables
  • running local open models inside coding agent harnesses
  • when is LoRA finetuning worth it vs base model plus prompting
  • small model ceiling for agentic tool use and long context

Measured heat

no measured readings yet โ€” the hourly heat pass fills this in

How the heat travelled

09-28 21:52โญ origin directly observedSearching for 3.8 35B: Qwen3.6-35B-A3B (Testing 5 Finetunes vs. Base)
returnity on r/LocalLLaMA
โ€”
09-28 21:52amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wss436
returnity
peak 61 ยท 61 comments ยท 100% of case engagement
09-28 23:20our radar first saw it ยท +1.5hdiscovery anchor: reddit.post.1wss436โ€”

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญSearching for 3.8 35B: Qwen3.6-35B-A3B (Testing 5 Finetunes vs. Base)
LocalLLaMA
returnity6363

Interpretation history

Decision trace