Skip to content
View original post on X: Rohan PaulX· 60/100AI score60/100

Xiaomi's MiMo-V2.6 paper details scaling RL for self-improving coding agents

AISummary

Xiaomi's MiMo-V2.6 paper says agents now build tasks, audit tests, grade answers and detect cheating during RL training, with humans setting the budget and rules.

A grader agent that rewards cleaner patches over reward-hacking fixes is credited with stopping drift toward longer runs and workarounds such as swallowed exceptions.

MiMo-V2.6-Pro's DeepSWE score rose from 58.4 to 72.6 over $2.6M of RL compute and was still climbing when training stopped.

Post on XView on X
Rohan PaulVerified on X
@rohanpaul_ai

The paper for MiMo-V2.6 by Xiaom is out.

In MiMo-V2.6, AI runs much of its own training loop: agents build the tasks, audit the tests, grade the answers and hunt for cheats. Humans set the budget and the rules.

shows that agent models kept improving with more RL compute by scaling batch size, task and harness variety, and grading effort together.

Scaling RL for coding agents is hard because pass/fail tests can't tell a clean fix from a hacky fix. Agents also learn to game environments, for example by downloading the published fix.

A grader agent compared passing patches in each group and moved reward to the cleaner ones. Without it, agents drifted toward longer runs and workarounds like swallowed exceptions.

MiMo-V2.6-Pro’s DeepSWE score rose from 58.4 to 72.6 over $2.6M of RL and was still climbing when training stopped.

– arxiv. org/abs/2610.11959

Title: "MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement"

Source: Rohan Paul · x.comPublished