Skip to content
Read the original: Prime Intellect· Published 34/100AI score34/100

Qwen3.6 reward rises 2.8x via GRPO on Hosted Training

Original titleAfter ~100 GRPO steps on Hosted Training, Qwen3.6's reward on held-out problems rose from 0.127 to 0.361, a 2.8x gain.

AISummary

Prime Intellect reports that after about 100 GRPO steps on Hosted Training, Qwen3.6's reward on held-out problems rose from 0.127 to 0.361, a 2.8x gain. Qwen3.5, trained the same way, reached 0.356, suggesting the method works across model families.

Both post-trained models finished well ahead of other open models and narrowed the gap to Claude Opus 4.8, with Qwen3.6 activating only 3B parameters per token.

Read the original x.com

Source: Prime Intellect · x.comPublished · added here