2026-10-11 18:03 UTC

Xiaomi's MiMo v2.6 launch introduces Pro and Flash variants alongside a published 9B Qwen distillation, expanding model choices for coding-agent and local-inference deployments without yet establishing comparative performance.

state: resolvedheat: mediumuncertainty: mediumconvergesscott: mediummodel-releases open-models local-inference coding-agentsXiaomi MiMo
Surfaced 2026-09-22T00:23:05Z β€” The linked Xiaomi page presents MiMo v2.6; accompanying Reddit observations report Pro and Flash releases and link Xiaomi's MiMo-V2.6-Distil β€” The XiaomiMiMo Hugging Face artifact turns the 9B distillation from an unverified launch claim into a concrete local-inference candidate. Rapid cross-community spread and early output testing justify immediate attention, although comparative capability and deployment quality remain unvalidated.

What is this?

On September 22, 2026, Xiaomi open-sourced the MiMo-V2.6 series: two natively omnimodal models, Pro and Flash, plus a MIT-licensed MiMo-V2.6-Distill-Qwen-9B, a technical report, 7K+ RL task environments, and an end-to-end RL framework. Xiaomi claims Pro scores 46 on the Artificial Analysis Intelligence Index β€” the strongest open-source model, ahead of Kimi K3 and Qwen3.8 Max but behind closed models β€” with TechNode reporting ~$2.62M (Pro) and ~$850K (Flash) RL training costs over 30 RL steps in under six days. Community reception is mixed: early hands-on reports cite vLLM serving defects, tool-use problems, and skepticism that benchmark gains translate to real coding work. A teased MiMo-V3 built on a newly published HySparse2 sparse architecture suggests the line is still accelerating.

Why it matters to Scott

Independent hands-on evidence now converges with Scott's canon: the 'benchmaxxed' accusation and senior-SWE testing finding benchmarks don't translate to real coding are exactly his Benchmarking the Wrong Unit / Model-Plus-Harness argument arriving from outside, and the capability-audit standard (representative production data, not demo conditions) is the right lens for matching the RL checkpoint to the advertised scores. It also bears on what he actually runs: the MIT 9B Qwen distill is a concrete eval candidate for the gamepc/Ollama zoo, the M5 Ultra throughput and dual-Spark tooling reports are exactly the hardware-aware-local-inference evidence his assessment called unresolved, and the vLLM empty-response/hidden-output-cap defects land in the same tool-call parsing failure family his `ask` harness (json-repair, multi-format parsing) already defends against β€” worth testing MiMo-Flash against his own tool-call layer, but the evidence still doesn't justify changing model selections.
ip:concept.benchmarking-the-wrong-unitip:concept.model-plus-harness-benchmark-unitip:concept.capability-auditdev:concept.hardware-aware-local-inferencedev:concept.multi-format-tool-call-parsingradar:xiaomi-mimo-live-training-dashboardradar:person.xiaomi-mimoradar:concept.open-modelsradar:concept.local-inferenceradar:concept.vllmradar:vllm-silent-tool-parser-failures
queries asked of Scott's wikis
  • open-weights frontier strategy: does a Xiaomi-level entrant change lab risk or leverage
  • local inference economics: memory footprint and hardware requirements for large open models vs small distills
  • coding-agent harness compatibility: vLLM tool-use and streaming defects as harness-side failure modes
  • benchmark vs real-world coding evaluation: benchmaxxing and agent task reliability
  • distillation and RL recipe patterns: SFT-on-teacher-data distills, RL environments as moat
  • sparse attention architectures: HySparse2 and prior sparse/efficient-architecture positions

Measured heat

no measured readings yet β€” the hourly heat pass fills this in

How the heat travelled

09-21 21:23 (minted)⭐ origin echo-reconstructedThe linked Xiaomi page presents MiMo v2.6; accompanying Reddit observations report Pro and Flash releases and link Xiaomi's MiMo-V2.6-Distil
Xiaomi MiMo on blog (echo) Β· attributed from reddit.post.1wmoxgz, hn.story.49792730, reddit.post.1wmp2di Β· published time unknown
β€”
09-21 20:12first on hacker news Β· published Β· lag ?Xiaomi MiMo v2.6
volf_
β€”
09-21 20:46first on r/singularity Β· published Β· lag ?Mimo v 2.6 pro and flash released
No-Selection2972
β€”
09-21 20:51first on r/LocalLLaMA Β· published Β· lag ?XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
VoiceApprehensive893
β€”
09-22 07:56first on r/MachineLearning Β· published Β· lag ?Xiaomi releases MiMo-V2.6: "Frontier intelligence, all the modalities, built in public." [N]
we_are_mammals
β€”
09-24 15:53first on r/artificial Β· published Β· lag ?I switched my personal agent from DeepSeek V4.1 Flash to MiMo V2.6 Pro
buffering_112
β€”
09-21 20:12amplified on hacker news πŸ‘‘hn.story.49792730
volf_
peak 1120 Β· 461 comments Β· 59% of case engagement
09-21 20:46amplified on r/singularityreddit.post.1wmoxgz
No-Selection2972
peak 45 Β· 9 comments Β· 1% of case engagement
09-21 20:51amplified on r/LocalLLaMAreddit.post.1wmp2di
VoiceApprehensive893
peak 289 Β· 71 comments Β· 7% of case engagement
09-21 23:50amplified on r/LocalLLaMAreddit.post.1wmtjqu
Beamsters
peak 188 Β· 39 comments Β· 4% of case engagement
09-22 04:02amplified on hacker newshn.story.49796660
theanonymousone
peak 162 Β· 65 comments Β· 8% of case engagement
09-22 07:07amplified on r/LocalLLaMAreddit.post.1wn2dee
anovers
peak 32 Β· 13 comments Β· 1% of case engagement
12 more amplifiers in ainews.case_chain
09-21 21:20our radar first saw it Β· lag ?discovery anchor: reddit.post.1wmoxgzβ€”
09-22 00:23reached heat=high Β· lag ? Β· via ledgerβ€”β€”

Evidence (19) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditMimo v 2.6 pro and flash released
singularity
No-Selection2972459
🟧 hnXiaomi MiMo v2.6volf_1123457
🟠 redditXiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
LocalLLaMA
VoiceApprehensive89328979
🟧 echo.blog ⭐The linked Xiaomi page presents MiMo v2.6; accompanying Reddit observations report Pro and Flash releases and link Xiaomi's MiMo-V2.6-DistilXiaomi MiMoβ€”β€”
🟠 redditMiMo-V2.6 distilled themselves into Qwen 9B!
LocalLLaMA
Beamsters18741
🟧 hnMiMo-v2.6-Pro: Intelligence, Performance and Price Analysistheanonymousone16467
🟠 redditXiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
LocalLLaMA
anovers3314
🟠 redditXiaomi releases MiMo-V2.6: "Frontier intelligence, all the modalities, built in public." [N]
MachineLearning
we_are_mammals9017
🟠 redditXiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B · Hugging Face
LocalLLaMA
Aggravating-Push-20711531
🟠 redditAbout Mimo 2.6 Architecture
LocalLLaMA
BagComprehensive797426
🟠 redditMimo v2.6-Flash-RL Dual-Spark Recipes/experiences?
LocalLLaMA
IamFondOfHugeBoobies612
🟠 redditMimo2.6-Flash on M5U early results
LocalLLaMA
bakawolf1231416
🟠 redditMiMo-V2.6-Flash on vLLM: fixes for "empty responses" with thinking + tools, and a hidden 2,048-token output cap
LocalLLaMA
mamolengo2115
🟠 redditMiMo-V3 is getting a new architecture. The core of it, HySparse2, is out today.
LocalLLaMA
Recoil4232949
🟠 redditMiMo-V2.6 (both Pro and Flash) is a benchmaxxed scam
LocalLLaMA
crusaderky7878
🟧 hnXiaomi MiMo v2.6 Pro: New best open-weight LLM with simple design
Retrieved article excerpt

Open article Β· Retrieved 2026-09-23T15:28:56.400335+00:00

Xiaomi’s new MiMo-V2.6 Pro is β€œsimply” the best (for now). Despite its simple architecture design it’s currently No.1 in the open-weight benchmarks (weighted average).

With β€œsimple,” I mean a classic [Grouped Query Attention (GQA)](https://sebastianraschka.com/llm-architecture-gallery/gqa/) with [Sliding Window Attention (SWA)](https://sebastianraschka.com/llm-architecture-gallery/swa/) at a tiny 128-token window size.

So, that underlines one of the points I’ve been trying to make in recent months: most of the progress still comes from the data and post-training recipe improvements. Fancy [attention variants](https://magazine.sebastianraschka.com/p/visual-attention-variants) are just mostly efficiency tweaks.

What are some of the training data improvements and recipe improvements? The MiMo team shared a pretty detailed [technical report](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL/blob/main/MiMo_V2_6_technical_report.pdf). Lots to carefully digest there, but in short, there are a few things that stood out:

1. An increase in agent tasks; also training across different harnesses (the average DeepSWE pass@1 accuracy on held-out harnesses improved from approximately 50% -> 66%).
2. Better reward signals: they replaced a simple correctness verifier with an agentic grader that looks at the execution traces as well.
3. Large [RL](https://magazine.sebastianraschka.com/p/the-state-of-llm-reasoning-model-training) batches (1,568 prompts Γ— 16 rollouts = 25,088 trajectories) and 2.7–3.7 billion training tokens per update (unclear, though, what the predecessor used).

Composite figure comparing MiMo-V2.6 Pro and DeepSeek V4-Pro architectures, Artificial Analysis Intelligence Index scores, and output speeds

MiMo-V2.6 Pro and DeepSeek V4-Pro architectures, with release-time Artificial Analysis Intelligence Index and output-speed comparisons.

Source: website version of my [Substack note](https://substack.com/@rasbt/note/c-343108099).
ModelForge10
🟧 hnMiMo-V3 is getting a new architecture. The core of it, HySparse2, is out todayksec10
🟠 redditMiMo 2.6 Pro: Reducing overthinking and second-guessing
LocalLLaMA
PilgrimofHaqq2410
🟠 redditI switched my personal agent from DeepSeek V4.1 Flash to MiMo V2.6 Pro
artificial
buffering_11273

Interpretation history

Decision trace