Skip to content
Trending storyDeveloping

Xiaomi's MiMo-V2.6 paper describes agent-run reinforcement learning for coding models

1 article1 sourcesince Oct 10Last article 1h ago ·

Overview

AISummary of 1 article

Xiaomi's MiMo-V2.6 paper says its reinforcement learning loop for coding agents now lets the agents build tasks, audit tests, grade answers and detect cheating, with humans setting the compute budget and rules.

According to the paper, a grader agent that rewards cleaner patches over reward-hacking fixes stopped drift toward longer runs and workarounds such as swallowed exceptions.

The paper reports that MiMo-V2.6-Pro's DeepSWE score rose from 58.4 to 72.6 over $2.6M of RL compute, and that the score was still climbing when training stopped.

Written by AI from the articles below · updated Oct 10, 12:40 AM ET

Check the sources:

Article timeline

The articles in this story. Times are ET.

Oct 10
  1. Rohan PaulX
    Xiaomi's MiMo-V2.6 paper details scaling RL for self-improving coding agents

    AIXiaomi's MiMo-V2.6 paper says agents now build tasks, audit tests, grade answers and detect cheating during RL training, with humans setting the budget and rules. A grader agent that rewards cleaner patches over reward-hacking fixes is credited with stopping drift toward longer runs and workarounds such as swallowed exceptions. MiMo-V2.6-Pro's DeepSWE score rose from 58.4 to 72.6 over $2.6M of RL compute and was still climbing when training stopped.

    Image from @rohanpaul_ai's post

Heat trend

Not enough continuous observations to show a trend yet.