Xiaomi says its MiMo-V2.6 update diagnoses and fixes tool-call repetition โ repeated identical tool calls burning context and stalling agent tasks in MiMo Desktop, MiMo Code and OpenCode โ attributing it to a 'reward blind spot' in scaled RL; confirmation that the patch ends the stalling in real agent workflows would establish post-release reward-blind-spot patching as a recognized failure mode of RL-trained open coding models.
state: seedheat: lowuncertainty: mediumconvergesscott: highagent-harnesses rl-reward-design open-model-updates local-inferenceXiaomi
What is this?
Xiaomi released and open-sourced the MiMo-V2.6 series on 2026-09-22, framing it on its own site as a step on the 'RSI path: scaling RL compute,' backed by an unusually public live RL training dashboard (mimo.xiaomi.com/rl) and a technical report on arXiv. The series is positioned for long-running agent workflows โ MiMo Desktop, MiMo Code, third-party harnesses like OpenCode โ with the product page touting 'Model x Harness Synergy' and 1M-token context; third-party coverage (InfoWorld, myclaw.ai, r/LocalLLaMA) treats it as a serious open contender for scaled agent work while cautioning that launch results are mostly vendor-reported with limited independent testing. The case's specific claim โ that a V2.6 update diagnoses tool-call repetition (identical calls burning context and stalling agent loops) as a 'reward blind spot' in the scaled RL run and fixes it โ comes from Xiaomi's own announcement captured in the case evidence; the supplied web results confirm the release and the RL positioning but do not independently corroborate the repetition diagnosis or that the patch resolves it in real workflows.
Why it matters to Scott
Xiaomi independently landing on a 'reward blind spot' diagnosis is a dated, first-party vendor receipt for the Goodhart/specification-gaming position already in Scott's canon โ scaled RL shipping a degenerate tool-call loop that only surfaced inside real harnesses โ and the failure itself (identical calls burning context until the task stalls) sits squarely in his Five-Surface Loop Anatomy / Handover Notes stuck-detection territory, with OpenCode explicitly named in the patch, a harness he actively runs. The live tension worth tracking: Xiaomi is patching model-side via RL while his whole playbook argues repetition loops need harness-side external detection โ and if blind spots prove to be a class rather than a one-off bug, that strengthens his guardrail argument even as vendor patches land.
ip:concept.specification-gamingip:framework.five-surface-loop-anatomyip:source.handover-notes-for-robots-ebookip:concept.evaluation-driven-developmentdev:technology.opencodedev:project.askradar:poolside-laguna-s-2-1-context-fixradar:codex-july-looping-regressionradar:concept.reward-hackingradar:concept.coding-modelsradar:concept.open-models
queries asked of Scott's wikis
- reward hacking and reward blind spots in agentic RL for coding agents
- tool-call repetition loops: harness-side detection and guardrails
- post-release update cadence for open coding models
- model-harness co-optimization: vendor-first agents vs third-party harnesses
- long-running agent context management and token burn
- open weights strategy: can open models hold agent-workload parity
Measured heat
now 0 pts/hpeak 18 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 321h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p46 vs 1188 stories at the 168h mark (now 321h old) โ ahead of aipass-false-success-fixes (1.1x), behind agentgit-accountless-agent-handoffs (0.9x)
Evidence (2) โ โญ canonical anchor
Interpretation history
2026-09-28T07:35:27Z
grounded: converges/high โ Xiaomi independently landing on a 'reward blind spot' diagnosis is a dated, first-party vendor receipt for the Goodhart/specification-gaming position already in
2026-09-28T07:26:39Z
case created โ First-party vendor diagnosis of a specific RL reward blind spot producing a concrete agent-loop failure, with the X announcement visible through the LocalLLaMA echo and no open case covering it.
Decision trace
- 09-28 17:35groundXiaomi independently landing on a 'reward blind spot' diagnosis is a dated, first-party vendor receipt for the Goodhart/specification-gaming position already in Scott's canon โ scaled R
- 09-28 17:26createFirst-party vendor diagnosis of a specific RL reward blind spot producing a concrete agent-loop failure, with the X announcement visible through the LocalLLaMA echo and no open case covering it.