Beating AI News Flash: The Xiaomi MiMo team has reviewed the issue of repeated tool calls after the launch of MiMo-V2.6. The model sometimes repeatedly calls the same or highly similar tools, continuously consuming context while making no progress on the task. In OpenCode, the proportion of replies with repeated tool calls for Flash and Pro once reached 1.02% and 0.54%, respectively.
The problem lies in the reward design of reinforcement learning. Training mainly rewarded "whether the task was ultimately done correctly," but did not sufficiently penalize inefficient behavior during the process. The original rule only imposed a penalty when a single round of tool calls exceeded 32 times, and repeated calls below 32 times were not penalized at all. As the scale of RL expanded, this bad habit was instead continuously reinforced. After replaying training checkpoints, Xiaomi found that in Flash, the proportion of abnormal samples with more than 10 calls in a single round increased from 11.1% at step 0 to 24.6% at step 20.
The most direct solution was to lower the penalty threshold from 32 times to 8 times and rerun about 20 MixRL steps, but this was expected to cost $2.31 million. Xiaomi ultimately trained only one RL teacher specifically to correct repeated calls, running 12 steps with about 7,000 samples, and then merged this capability back into Pro and Flash through MOPD. The entire round of fixes cost about $90,000, only about 4% of the full retraining plan, and other major benchmarks remained basically unchanged.
The fixed MiMo-V2.6-Pro-MOPD and Flash-MOPD weights have already been open-sourced, and the API has also been switched to the new version, with the call names unchanged. Xiaomi will also reset the remaining quota for MiMo Desktop users in the current cycle.

