2026-10-11 18:04 UTC

Independent reproduction will determine whether supervised fine-tuning with Qwen3’s default chat template can suppress thinking mode while leaving training loss and basic evaluations apparently normal.

state: expiredheat: lowuncertainty: highnovelscott: lowqwen fine-tuning reasoning-modelsQwenMeldh LLC

What is this?

The case alleges that a Qwen3-8B supervised/LoRA fine-tuning run on non-thinking data lost its thinking behavior while ordinary training metrics and basic evaluations remained normal, reportedly involving Meldh LLC. The supplied web snippets establish that Qwen3-family thinking behavior is mediated partly through chat-template controls such as `enable_thinking`, with inconsistent support across models and runtimes. They do not independently document the reported fine-tuning experiment, identify Meldh LLC’s role, or establish that the default template can silently suppress thinking after training, so the central claim remains unverified.

Why it matters to Scott

No intersection was found in Scott’s wikis or existing radar pages. The unverified result may be relevant to reasoning-model fine-tuning practice, but the supplied material does not connect it to a position or active project of Scott’s.
queries asked of Scott's wikis
  • fine-tuning silent capability regressions
  • reasoning-model evaluation beyond training loss
  • chat templates as model behavior control
  • LoRA catastrophic forgetting reasoning
  • independent reproduction of fine-tuning results
  • thinking-mode evaluation harnesses

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditDuring fine-tuning of Qwen3-8B, one build lost its thinking mode and the training metrics never noticed.
LocalLLaMA
MeldhLLC00
🟧 echo.blog ⭐The original report describes experiments showing that standard LoRA fine-tuning of Qwen3-8B on non-thinking data improved non-reasoning accMaruf Bepary——

Interpretation history

Decision trace