2026-10-11 17:09 UTC

reasoning-models

band: coolmomentum: stable score: 0.268
temperature history

Episodes (2)

The paper’s authors claim increasing expert activation only in the later layers of Qwen sparse-MoE models reduces reasoning-token use by about 8.5% without retraining or material quality loss, potentially lowering inference cost through a runtime-only change.
watchingconvergesscott: medium
Independent reproduction will determine whether supervised fine-tuning with Qwen3’s default chat template can suppress thinking mode while leaving training loss and basic evaluations apparently normal.
expirednovelscott: low