2026-10-11 17:12 UTC

Independent reproduction will determine whether the reported train-inference mismatch in open-weight MoE reinforcement-learning stacks is a widespread failure mode that materially undermines reproducibility and deployed-model quality.

state: expiredheat: lowuncertainty: highconvergesscott: mediumopen-models reinforcement-learning moe-models training-inferencekiddyboots216

What is this?

Researchers have documented a rollout/train-inference mismatch in LLM reinforcement-learning systems when separate inference and training stacks produce different token probabilities despite sharing model weights, effectively making nominally on-policy RL off-policy. The supplied results report that this can destabilize or collapse training, with MoE models particularly exposed because small logit differences can activate different experts; proposed mitigations include truncated importance sampling and bitwise-consistent training/inference stacks. The snippets include experiments on Qwen3 dense and MoE models, but they do not establish the identity or role of “kiddyboots216,” and the claimed independent reproduction is only partially evidenced by separate reports rather than a clearly described reproduction of the cited Wordle experiment.

Why it matters to Scott

The reported mismatch extends Scott’s non-determinism and model-plus-harness positions into RL training: identical weights are insufficient when training and serving stacks produce behaviorally different policies, making reproducible evaluation a systems property. This is actionable for open-weight RL pipeline design, but independent reproduction and the breadth of the MoE failure mode remain unestablished, and the hits show no active Scott project using this stack.
ip:concept.non-determinismip:concept.model-plus-harness-benchmark-unitip:concept.evaluation-driven-developmentradar:concept.reinforcement-learningradar:concept.vllmradar:concept.llm-inferenceradar:concept.mixture-of-expertsradar:concept.open-models
queries asked of Scott's wikis
  • rollout-training mismatch in on-policy LLM RL
  • shared kernels and bitwise-consistent training inference
  • MoE router instability across serving and training stacks
  • reproducibility of open-weight RL pipelines
  • vLLM training harness policy divergence
  • systems-level failure modes in agent reinforcement learning

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnTrain-infer mismatch for Open-weight MoE RL in Open-source codekiddyboots21610
🟧 echo.blog ⭐The page is the original technical article. It says, “We train Qwen3.6-35B-A3B to play Wordle” and that eliminating train-infer mismatch impAshwinee Panda——

Interpretation history

Decision trace