Skip to content
Read the original: Alexander Doria· Dorialexander·Published AI score38/100

Since I went into this release: *It's a smaller selection of 989 envs used to RL a 9B distilled model, not the big MiMo.

Original titleSince I went into this release:

AISummary

Rewards are not self-contained: general part need to set up a judge and webdev rely on their own grader service+vlm. *Most important part is inside general/envs directory (+ docker), not the dataset displayed on hf: genuinely solid mix of real/simulated documents we rarely see in OSS.

Read the original x.com

Source: Alexander Doria · x.com