Xiaomi's MiMo team publicly live-streamed the reinforcement-learning post-training of its next models: a dashboard at mimo.xiaomi.com/rl ('Open is what we value') streamed reward curves, running compute cost, task-mix telemetry, and mid-run benchmark scores for MiMo-V2.6-Pro and -Flash starting ~September 15, 2026, and per Xiaomi's own docs both runs completed 30 steps in ~6 days at reported costs of roughly $850k (Flash) and $2.62M (Pro). Team lead Fuli Luo framed it as a test of 'how far RL can scale' (fully-async runs, ~2B tokens/step, 1,568 prompts x 16 rollouts), and the models are now listed in Xiaomi's docs as API-callable (mimo-v2.6-pro/flash/pro-ultraspeed, with Cline integration), with Hugging Face repos sighted and Sebastian Raschka publishing independent architecture/training notes ranking MiMo-V2.6 Pro No.1 among open-weight models on weighted-average benchmarks. The supplied material is still thin or conflicting on specifics that matter: parameter counts are secondhand (~1.02T/42B-active Pro vs 309B/15B Flash per one Substack, against other community figures), whether checkpoints are genuinely downloadable open weights with usable licenses versus API-only is not pinned down, and the community-cited DeepSWE scores conflict with dashboard-derived step-30 figures (65-73%). The separate open MiMo-V3 run with a new HySparse2 architecture is not addressed by these snippets.
Xiaomi's open RL run is the corporate sibling of the radar's Liang open-training episode, but Scott's canon holds no position on training-run transparency — the dashboard is his own W&B-style training-telemetry pattern reappearing on someone else's run, and the one Scott-bearing payload, a potential top-open-weight candidate for his Model-Barbell/LiteLLM routing, stays gated on unverified weights, licenses and independent harness evals (Raschka's No.1 ranking aggregates self-reported dashboard benchmarks, precisely the producer-run scoring his Model-Plus-Harness Benchmark Unit distrusts). Upgrade paths: verified downloadable 2.6 checkpoints under a usable license, community harness benchmarking, or usable artifacts from the live V3 HySparse2 open run.
ip:concept.model-barbellip:concept.model-plus-harness-benchmark-unitdev:technology.litellmradar:percy-liang-open-535b-training-runradar:concept.model-trainingradar:concept.open-modelsradar:concept.post-trainingradar:concept.observabilityradar:open-moe-train-inference-mismatch
queries asked of Scott's wikis
- live training-run telemetry observability dashboards wandb
- open-weight model selection coding agent backbone local inference
- agentic RL environments rollout harness reward grader post-training
- eval harness comparability held-out benchmark contamination
- training-run transparency open frontier lab practice
- MoE active parameters serving cost price-performance
2026-09-27T19:25:29Z
The reported release of MiMo-V2.6-Flash-MOPD weights — the output of the very run the dashboard livestreamed — closes the arc from live-telemetry spectacle to produced artifact, settling the case's question: live post-training visibility is a real, first-party Xiaomi practice that ended in released models. The episode resolves as absorbed; the forward substance (independent harness evals of the released 2.6 weights, the V3 HySparse2 open run) belongs to fresh episodes, not this one held open on sequel promise.
2026-09-27T18:24:42Z
evidence attached: reddit.post.1wrq71o — The released MiMo-V2.6-Flash-MOPD weights appear to be the output of the live post-training run the dashboard case tracks, completing that episode.
2026-09-26T04:34:02Z
grounded: novel/low — Xiaomi's open RL run is the corporate sibling of the radar's Liang open-training episode, but Scott's canon holds no position on training-run transparency — the
2026-09-26T04:25:32Z
Attention has collapsed (0.17 pts/h vs a 555/h peak; the new HN item sits at 1 point, 0 comments), but the substantive frontier moved: the MiMo lead's announcement of the open MiMo-V3 run with a new HySparse2 architecture turns the case from a one-off 2.6 dashboard spectacle into an ongoing Xiaomi open-training arc. A first-party continuation announcement is a second, independent evidence line behind the multi-platform eyewitness reports, so the phenomenon graduates from rumored to corroborated — while the 2.6 telemetry, weights and benchmark claims themselves remain unverified — and heat cools high→medium rather than to low because the V3 thread is live.
2026-09-26T04:23:19Z
evidence attached: hn.story.49853136 — MiMo lead announcing the open MiMo-V3 run's new HySparse2 architecture is first-party detail the open-training-visibility case must carry.
2026-09-21T21:53:43Z
A second poster links the same Pro-RL repository, broadening circulation without verifying model availability, repository contents or live training telemetry. This is repeated coverage rather than substantive corroboration; high heat remains justified by the episode’s broad Reddit/Hacker News reach and recently emerging artifact discussion, not stronger capability evidence.
2026-09-21T21:22:54Z
evidence attached: reddit.post.1wmo81b — shared external link with case evidence
2026-09-21T20:58:07Z
Reported Hugging Face repositories for Flash-RL and Pro-RL extend the episode beyond dashboard spectatorship toward potentially inspectable model artifacts, without yet establishing downloadable weights or usable access. Strong cross-platform attention plus this expanding artifact periphery warrants high heat, although live-telemetry and capability claims remain unverified.
2026-09-21T20:57:49Z
evidence attached: reddit.post.1wmnw89, reddit.post.1wmntwv — The Flash-RL and Pro-RL repository sightings are companion artifacts of the already tracked MiMo 2.6 post-training episode, not grounds for a second family-level case.
2026-09-17T14:42:18Z
The surfaced Xiaomi-domain URL makes the dashboard claim easier to investigate, but supplies no inspected telemetry or new capability evidence. Remaining discussion amplifies the transparency narrative without changing its practical significance for Scott.
2026-09-17T07:31:01Z
A new commenter supplies specific purported DeepSWE scores and training steps, making the telemetry claim more concrete but not independently verified. The frontier-performance comparison remains unsupported without evaluation methodology or reproducible results, so this does not yet change Scott’s model-selection decisions.
2026-09-16T22:40:44Z
The added post specifies PRO/FLASH variants and alleges a $10-per-second burn rate, but remains another account of the same dashboard rather than independent validation of live training telemetry. Discussion broadens awareness without establishing a model-selection or evaluation consequence for Scott.
2026-09-16T22:22:10Z
evidence attached: reddit.post.1wibfhj — The post independently corroborates Xiaomi's public live post-training dashboard and adds evidence about its visible burn rate.
2026-09-16T21:28:59Z
grounded: novel/low — The reported dashboard is at most another example of inspectable training metrics, adjacent to Scott’s Observability concept and his use of Weights & Biases; th
2026-09-16T21:23:28Z
case created — Three discussions point to one concrete first-party dashboard, although they do not independently validate its contents.