Xiaomi's model reaches frontier performance in just 30 RL steps
Original titleRespect to Xiaomi for pacing the frontier using only 30 steps
AISummary
A post praises Xiaomi for matching frontier-level results after only 30 reinforcement learning steps. A quoted reply notes the model improved dramatically in those 30 steps, with no sign of a plateau, and wonders why training stopped there.
Source: wh · x.comPublished · added here