Independent reproduction will determine whether the reported train-inference mismatch in open-weight MoE reinforcement-learning stacks is a widespread failure mode that materially undermines reproducibility and deployed-model quality.
state: expiredheat: lowuncertainty: highconvergesscott: mediumopen-models reinforcement-learning moe-models training-inferencekiddyboots216
What is this?
Researchers have documented a rollout/train-inference mismatch in LLM reinforcement-learning systems when separate inference and training stacks produce different token probabilities despite sharing model weights, effectively making nominally on-policy RL off-policy. The supplied results report that this can destabilize or collapse training, with MoE models particularly exposed because small logit differences can activate different experts; proposed mitigations include truncated importance sampling and bitwise-consistent training/inference stacks. The snippets include experiments on Qwen3 dense and MoE models, but they do not establish the identity or role of “kiddyboots216,” and the claimed independent reproduction is only partially evidenced by separate reports rather than a clearly described reproduction of the cited Wordle experiment.
Why it matters to Scott
The reported mismatch extends Scott’s non-determinism and model-plus-harness positions into RL training: identical weights are insufficient when training and serving stacks produce behaviorally different policies, making reproducible evaluation a systems property. This is actionable for open-weight RL pipeline design, but independent reproduction and the breadth of the MoE failure mode remain unestablished, and the hits show no active Scott project using this stack.
ip:concept.non-determinismip:concept.model-plus-harness-benchmark-unitip:concept.evaluation-driven-developmentradar:concept.reinforcement-learningradar:concept.vllmradar:concept.llm-inferenceradar:concept.mixture-of-expertsradar:concept.open-models
queries asked of Scott's wikis
- rollout-training mismatch in on-policy LLM RL
- shared kernels and bitwise-consistent training inference
- MoE router instability across serving and training stacks
- reproducibility of open-weight RL pipelines
- vLLM training harness policy divergence
- systems-level failure modes in agent reinforcement learning
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-19T20:39:47Z
No independent reproduction or implementation surfaced within the case’s freshness horizon, so the reported MoE train-inference mismatch remains a technically concrete but isolated result rather than an active developing episode.
2026-08-17T19:45:58Z
No independent reproduction, implementation, or consequential entrant has appeared; the original single-team result remains technically concrete but does not yet establish a widespread MoE RL failure mode.
2026-08-17T19:39:44Z
grounded: converges/medium — The reported mismatch extends Scott’s non-determinism and model-plus-harness positions into RL training: identical weights are insufficient when training and se
2026-08-17T19:36:35Z
origin walked (codex/luna, conf 0.97): anchor hn.story.49336172 -> echo.blog.2922b1b027 by Ashwinee Panda
2026-08-17T19:34:57Z
case created — Concrete engineering report on a specific failure mode in open MoE RL stacks, but low engagement and no independent corroboration yet.
Decision trace
- 08-20 06:39expireNo independent reproduction or implementation surfaced within the case’s freshness horizon, so the reported MoE train-inference mismatch remains a technically concrete but isolated result rather than
- 08-20 06:39alert_silentThe only change is a staleness trigger with unchanged engagement and no new evidence; there is no consequential delta to surface.
- 08-20 06:39alert_routeThe only change is a staleness trigger with unchanged engagement and no new evidence; there is no consequential delta to surface.
- 08-18 05:45repriceNo independent reproduction, implementation, or consequential entrant has appeared; the original single-team result remains technically concrete but does not yet establish a widespread MoE RL failure
- 08-18 05:45alert_silentThis look contains only an unchanged reobservation of the existing report, with no new consequential evidence for Scott; it can wait for independent reproduction or broader implementation evidence.
- 08-18 05:45alert_routeThis look contains only an unchanged reobservation of the existing report, with no new consequential evidence for Scott; it can wait for independent reproduction or broader implementation evidence.
- 08-18 05:40alert_shadowA new original technical report provides a concrete implementation and experiment in which aligning training and serving behavior raised Qwen3.6-35B-A3B Wordle solve rate from 63.9% to 77.4%. The brea
- 08-18 05:40alert_routeA new original technical report provides a concrete implementation and experiment in which aligning training and serving behavior raised Qwen3.6-35B-A3B Wordle solve rate from 63.9% to 77.4%. The brea
- 08-18 05:39groundThe reported mismatch extends Scott’s non-determinism and model-plus-harness positions into RL training: identical weights are insufficient when training and serving stacks produce behaviorally differ
- 08-18 05:36promote_anchororigin walk conf 0.97
- 08-18 05:34createConcrete engineering report on a specific failure mode in open MoE RL stacks, but low engagement and no independent corroboration yet.