Xiaomi's MiMo 2.6 opens a live dashboard for its RL training runs
Xiaomi's MiMo team has published a public dashboard that streams reinforcement-learning metrics, benchmark scores and incident notes for mimo-v2.6-pro and mimo-v2.6-flash straight from the trainer's logs.
Xiaomi's MiMo team opened a public dashboard on 17 September that shows reinforcement-learning metrics for the mimo-v2.6-pro and mimo-v2.6-flash models, streamed live from the trainer's logs. The site is split into three views — overview, metrics and about — and includes a notice feed, per-metric charts, a benchmark table and the batch composition of every training step.
What the dashboard shows
The most unusual part is the notice feed, which openly lists failures that labs normally keep private. The pro run restarted at step 17 because of a GPU out-of-memory failure caused by expert load imbalance; the team adjusted the training parallelism strategy in response.
Another entry describes a network connectivity problem between the pro training cluster and the grader deployment. The run was restarted, and the cyber dataset was removed from the upcoming pro run after bad patterns appeared in the rollout logs. The flash run was restarted from step 15 because an infrastructure error on one dataset went undetected for roughly three hours. Both runs stopped at step 30: pro on 20 September after 5 days and 7 hours, flash on 19 September.
Numbers and benchmarks
The dashboard also publishes benchmark results at avg@3. On DeepSWE v1.1 with mini-swe-agent, pro scores 72.57 and flash 65.68; on an in-house Coding Bench, 65.43 and 62.87; on AutomationBench v1.0.6, 53.10 and 52.70.
Among the training metrics, dynsam/avg@n stands at 0.633 for pro (+0.011) and 0.644 for flash (−0.019). Mean context length is 137k tokens for pro and 148k for flash; tokens per step reach 3.43 billion and 3.7 billion, while a single step takes 6 hours 26 minutes and 3 hours 27 minutes respectively. In the last data posted, pro had accepted 664 samples against a target of 1,568 (pass rate 0.606), while flash had accepted 2,704.
Why it matters
Reinforcement-learning runs draw on thousands of datasets — the dashboard tracks 25 sources across code, cyber, visual, general and chat categories. Such operational detail is rarely made public. The MiMo dashboard shows that training a frontier model is a long and failure-prone process in which restarts and repeated debugging are the norm rather than the exception.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.