On September 22, 2026, Xiaomi open-sourced the MiMo-V2.6 series: two natively omnimodal models, Pro and Flash, plus a MIT-licensed MiMo-V2.6-Distill-Qwen-9B, a technical report, 7K+ RL task environments, and an end-to-end RL framework. Xiaomi claims Pro scores 46 on the Artificial Analysis Intelligence Index β the strongest open-source model, ahead of Kimi K3 and Qwen3.8 Max but behind closed models β with TechNode reporting ~$2.62M (Pro) and ~$850K (Flash) RL training costs over 30 RL steps in under six days. Community reception is mixed: early hands-on reports cite vLLM serving defects, tool-use problems, and skepticism that benchmark gains translate to real coding work. A teased MiMo-V3 built on a newly published HySparse2 sparse architecture suggests the line is still accelerating.
Independent hands-on evidence now converges with Scott's canon: the 'benchmaxxed' accusation and senior-SWE testing finding benchmarks don't translate to real coding are exactly his Benchmarking the Wrong Unit / Model-Plus-Harness argument arriving from outside, and the capability-audit standard (representative production data, not demo conditions) is the right lens for matching the RL checkpoint to the advertised scores. It also bears on what he actually runs: the MIT 9B Qwen distill is a concrete eval candidate for the gamepc/Ollama zoo, the M5 Ultra throughput and dual-Spark tooling reports are exactly the hardware-aware-local-inference evidence his assessment called unresolved, and the vLLM empty-response/hidden-output-cap defects land in the same tool-call parsing failure family his `ask` harness (json-repair, multi-format parsing) already defends against β worth testing MiMo-Flash against his own tool-call layer, but the evidence still doesn't justify changing model selections.
ip:concept.benchmarking-the-wrong-unitip:concept.model-plus-harness-benchmark-unitip:concept.capability-auditdev:concept.hardware-aware-local-inferencedev:concept.multi-format-tool-call-parsingradar:xiaomi-mimo-live-training-dashboardradar:person.xiaomi-mimoradar:concept.open-modelsradar:concept.local-inferenceradar:concept.vllmradar:vllm-silent-tool-parser-failures
queries asked of Scott's wikis
- open-weights frontier strategy: does a Xiaomi-level entrant change lab risk or leverage
- local inference economics: memory footprint and hardware requirements for large open models vs small distills
- coding-agent harness compatibility: vLLM tool-use and streaming defects as harness-side failure modes
- benchmark vs real-world coding evaluation: benchmaxxing and agent task reliability
- distillation and RL recipe patterns: SFT-on-teacher-data distills, RL environments as moat
- sparse attention architectures: HySparse2 and prior sparse/efficient-architecture positions
| source | object | author | score | comments |
| π reddit | Mimo v 2.6 pro and flash released singularity | No-Selection2972 | 45 | 9 |
| π§ hn | Xiaomi MiMo v2.6 | volf_ | 1123 | 457 |
| π reddit | XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B LocalLLaMA | VoiceApprehensive893 | 289 | 79 |
| π§ echo.blog β | The linked Xiaomi page presents MiMo v2.6; accompanying Reddit observations report Pro and Flash releases and link Xiaomi's MiMo-V2.6-Distil | Xiaomi MiMo | β | β |
| π reddit | MiMo-V2.6 distilled themselves into Qwen 9B! LocalLLaMA | Beamsters | 187 | 41 |
| π§ hn | MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis | theanonymousone | 164 | 67 |
| π reddit | XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B LocalLLaMA | anovers | 33 | 14 |
| π reddit | Xiaomi releases MiMo-V2.6: "Frontier intelligence, all the modalities, built in public." [N] MachineLearning | we_are_mammals | 90 | 17 |
| π reddit | XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B Β· Hugging Face LocalLLaMA | Aggravating-Push-207 | 115 | 31 |
| π reddit | About Mimo 2.6 Architecture LocalLLaMA | BagComprehensive79 | 74 | 26 |
| π reddit | Mimo v2.6-Flash-RL Dual-Spark Recipes/experiences? LocalLLaMA | IamFondOfHugeBoobies | 6 | 12 |
| π reddit | Mimo2.6-Flash on M5U early results LocalLLaMA | bakawolf123 | 14 | 16 |
| π reddit | MiMo-V2.6-Flash on vLLM: fixes for "empty responses" with thinking + tools, and a hidden 2,048-token output cap LocalLLaMA | mamolengo | 21 | 15 |
| π reddit | MiMo-V3 is getting a new architecture. The core of it, HySparse2, is out today. LocalLLaMA | Recoil42 | 329 | 49 |
| π reddit | MiMo-V2.6 (both Pro and Flash) is a benchmaxxed scam LocalLLaMA | crusaderky | 78 | 78 |
| π§ hn | Xiaomi MiMo v2.6 Pro: New best open-weight LLM with simple designRetrieved article excerptOpen article Β· Retrieved 2026-09-23T15:28:56.400335+00:00 Xiaomiβs new MiMo-V2.6 Pro is βsimplyβ the best (for now). Despite its simple architecture design itβs currently No.1 in the open-weight benchmarks (weighted average).
With βsimple,β I mean a classic [Grouped Query Attention (GQA)](https://sebastianraschka.com/llm-architecture-gallery/gqa/) with [Sliding Window Attention (SWA)](https://sebastianraschka.com/llm-architecture-gallery/swa/) at a tiny 128-token window size.
So, that underlines one of the points Iβve been trying to make in recent months: most of the progress still comes from the data and post-training recipe improvements. Fancy [attention variants](https://magazine.sebastianraschka.com/p/visual-attention-variants) are just mostly efficiency tweaks.
What are some of the training data improvements and recipe improvements? The MiMo team shared a pretty detailed [technical report](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL/blob/main/MiMo_V2_6_technical_report.pdf). Lots to carefully digest there, but in short, there are a few things that stood out:
1. An increase in agent tasks; also training across different harnesses (the average DeepSWE pass@1 accuracy on held-out harnesses improved from approximately 50% -> 66%).
2. Better reward signals: they replaced a simple correctness verifier with an agentic grader that looks at the execution traces as well.
3. Large [RL](https://magazine.sebastianraschka.com/p/the-state-of-llm-reasoning-model-training) batches (1,568 prompts Γ 16 rollouts = 25,088 trajectories) and 2.7β3.7 billion training tokens per update (unclear, though, what the predecessor used).
Composite figure comparing MiMo-V2.6 Pro and DeepSeek V4-Pro architectures, Artificial Analysis Intelligence Index scores, and output speeds
MiMo-V2.6 Pro and DeepSeek V4-Pro architectures, with release-time Artificial Analysis Intelligence Index and output-speed comparisons.
Source: website version of my [Substack note](https://substack.com/@rasbt/note/c-343108099). | ModelForge | 1 | 0 |
| π§ hn | MiMo-V3 is getting a new architecture. The core of it, HySparse2, is out today | ksec | 1 | 0 |
| π reddit | MiMo 2.6 Pro: Reducing overthinking and second-guessing LocalLLaMA | PilgrimofHaqq2 | 4 | 10 |
| π reddit | I switched my personal agent from DeepSeek V4.1 Flash to MiMo V2.6 Pro artificial | buffering_112 | 7 | 3 |
2026-09-24T17:54:06Z
The launch episode closes absorbed: the comparative-performance gap the hypothesis flagged is now filled β the benchmark lead is real but real coding-agent reliability is mixed and harness-dependent, the worst tool failures traced to fixable vLLM serving bugs, and cost efficiency is the strongest sustained claim. Velocity has collapsed to ~0.5% of peak with no expanding periphery, and the line's live momentum (V3/HySparse2) is its own story.
2026-09-24T16:38:17Z
evidence attached: reddit.post.1wp5ega β Independent user adoption with measured cost/quality comparison (AA 46 vs 39, 15-20% cheaper runs) filling the case's missing comparative-performance gap.
2026-09-24T02:44:37Z
magnitude valve eligible (multi-platform, top-decile engagement) and never alerted; deterministic escalation to deliver
2026-09-24T02:23:07Z
evidence attached: reddit.post.1wopeqg β Reproducible community evaluation of the released MiMo 2.6 Pro (-28% reasoning tokens, planted-premise test) is early independent quality evidence the launch case needs.
2026-09-24T01:01:21Z
No new comparative result changes the assessment; the latest activity mainly amplifies the already-established split between strong benchmark claims and weak early coding-agent reports. Heat remains high because discussion and implementations continue accelerating across several platforms, despite the absence of fresh material evidence.
2026-09-23T23:35:02Z
grounded: converges/medium β Independent hands-on evidence now converges with Scott's canon: the 'benchmaxxed' accusation and senior-SWE testing finding benchmarks don't translate to real c
2026-09-23T21:46:36Z
evidence attached: hn.story.49821132 β First-party announcement that MiMo-V3's new HySparse2 sparse architecture is out today extends the accelerating MiMo model-line episode.
2026-09-23T17:00:28Z
evidence attached: hn.story.49801771 β Raschka's architecture analysis is new evidence on the ongoing MiMo v2.6 release episode, noting the detailed technical report and training-recipe findings.
2026-09-23T16:27:31Z
evidence attached: reddit.post.1woa5d3 β Hands-on senior-SWE testing plus a corroborating top comment claim MiMo-V2.6's benchmark scores don't translate to real coding tasks β the first substantive counter-evidence on comparative performance for the launch case.
2026-09-23T15:28:49Z
evidence attached: reddit.post.1wo7mr6 β Xiaomi publishing the HySparse2 architecture underlying a teased MiMo-V3 is a material next-generation development in the accelerating MiMo release episode, echoed at 91 upvotes with an arXiv paper.
2026-09-22T20:23:58Z
evidence attached: reddit.post.1wnjbwn β Practical vLLM reports identify tool-use, streaming, and output-limit serving defects affecting MiMo-V2.6 deployments.
2026-09-22T15:22:43Z
evidence attached: reddit.post.1wnb5f5 β Independent early benchmarking shows MiMo v2.6 Flash running at roughly 49 tokens per second on an M5 Ultra, useful evidence about real local deployment performance.
2026-09-22T15:22:43Z
evidence attached: reddit.post.1wnbp77 β Early users report that MiMo v2.6 Flash has tooling problems on dual-Spark deployments, providing direct compatibility evidence for the model-release case.
2026-09-22T12:23:32Z
evidence attached: reddit.post.1wn6qwi β The discussion adds context to MiMo 2.6 by highlighting its conventional architecture and apparent emphasis on reinforcement learning.
2026-09-22T10:21:58Z
evidence attached: reddit.post.1wn59vj β shared external link with case evidence
2026-09-22T09:22:43Z
The release's periphery continues expanding across HN and Reddit, warranting sustained high attention, but the newly attached coverage supplies no inspectable comparative results or verified pricing. Reports that the 9B artifact is an SFT checkpoint, rather than the reportedly stronger RL version, introduce an evaluation caveatβnot evidence that the downloadable model delivers the advertised uplift.
2026-09-22T08:21:53Z
evidence attached: reddit.post.1wn36d4 β Independent Reddit coverage adds concrete release details, including the reported training cost and live progress dashboard.
2026-09-22T07:23:23Z
evidence attached: reddit.post.1wn2dee β The released 9B Qwen distillation is concrete independent evidence that MiMo 2.6 is expanding local open-model options.
2026-09-22T07:23:23Z
evidence attached: hn.story.49796660 β Independent analysis adds early intelligence, performance, and price evidence to the MiMo 2.6 release episode.
2026-09-22T00:23:05Z
The XiaomiMiMo Hugging Face artifact turns the 9B distillation from an unverified launch claim into a concrete local-inference candidate. Rapid cross-community spread and early output testing justify immediate attention, although comparative capability and deployment quality remain unvalidated.
2026-09-22T00:22:32Z
evidence attached: reddit.post.1wmtjqu β The released Qwen 9B distillation is a concrete artifact that materially strengthens the case for Xiaomi's MiMo 2.6 model family and local deployment options.
2026-09-21T21:29:33Z
grounded: known/low β The grounded MiMo V2.6 development is already tracked in radar:xiaomi-mimo-live-training-dashboard; the claimed subsequent launch and 9B distillation remain unv
2026-09-21T21:23:54Z
case created β The release and its small-model artifact form one launch episode distinct from the existing live-training-dashboard case; discussion volume indicates attention, not independent capability validation.