Xiaomi's MiMo-V2.6 report explains scaling RL along batch, environments, and grading
AIXiaomi's MiMo-V2.6 technical report argues that scaling reinforcement learning, not more pretraining data, is the main lever for frontier capability, along batch size, environment diversity, and grader strength. The article summarizes the report's methods, including groupwise agentic grading, a frozen-router fix for expert load collapse, and reward hacking defenses. It reports RL post-training costs of $2.6 million for MiMo-V2.6-Pro and $0.9 million for MiMo-V2.6-Flash.
Why it matters: The piece walks through the report's three-way RL scaling method, batch size, environments, and grading, with concrete failure modes and stabilization fixes useful to agentic RL practitioners.

