❯ Xiaomi’s MiMo-V2.6 reaches the middle of reinforcement learning, with the company showing training costs above $1.25 million
Training statusXiaomi said MiMo model lead Luo Fuli disclosed that MiMo-V2.6 is midway through reinforcement learning. A company training livestream showed cumulative spending above $1.25 million and described the model as approaching release.
Scaling workLuo said the team spent nearly six months exploring reinforcement-learning limits after open-sourcing MiMo-V2.5 in April. The current phase expands compute, environments and tools, and grader compute, with methods and engineering details due to be released over the coming weeks.
Proof pointPublic spending data makes the experiment easier to observe, but cost is not evidence of capability. Developers need post-release results on task success, tool reliability and unit inference cost before judging what the three scaling dimensions delivered.
▪ SIGNALThe livestream makes training more transparent; MiMo-V2.6 will still be judged by its released task performance.