Since I went into this release: *It's a smaller selection of 989 envs used to RL a 9B distilled model, not the big MiMo.
Original titleSince I went into this release:
AISummary
Rewards are not self-contained: general part need to set up a judge and webdev rely on their own grader service+vlm. *Most important part is inside general/envs directory (+ docker), not the dataset displayed on hf: genuinely solid mix of real/simulated documents we rarely see in OSS.
Source: Alexander Doria · x.com