Article

Xiaomi admits it MiMo-V2.6 will “repeat in place”:RL bad habits only get reinforced through training, 90000 USD Fix

Beating AI Flash News, the Xiaomi MiMo team reviewed the tool-call repetition issue after MiMo-V2.6 went live. The model sometimes repeatedly calls the same or highly similar tools, continuously consuming context without making progress on the task. In OpenCode, the shares of replies with repeated tool calls for Flash and Pro once reached 1.02%and 0.54%, respectively.

The problem lay in the reward design for RL. Training mainly rewarded whether the task was ultimately completed correctly, without sufficiently penalizing inefficient behavior during the process. The original rule imposed a penalty only when tool calls in a single round exceeded 32 calls; repeated calls below 32 calls incurred no deduction at all. As the scale of RL expanded, this bad habit was instead continually reinforced. After replaying training at checkpoint Xiaomi found that, in Flash, the proportion of anomalous samples with more than 10 calls in a single round rose from 0 at step 11.1%to 20 at step 24.6%.

The most direct approach was to lower the penalty threshold from 32 calls to 8 calls and rerun approximately 20 MixRL steps, but this was expected to cost 2310000 USD. Xiaomi ultimately trained a RL teacher specifically to correct repeated calls, ran it for 12 steps with approximately 7000 samples, and then used MOPD to merge this capability back into Pro and Flash. The entire fix cost approximately 90000 USD, only about 4%of the full retraining approach, while the other main benchmark remained essentially unchanged.

The fixed MiMo-V2.6-Pro-MOPD and Flash-MOPD weights are now available, and API has also switched to the new version; the call names remain unchanged. Xiaomi will also reset the remaining quota for MiMo Desktop users in the current cycle.

Original link https://m.theblockbeats.info/flash/369327