Sustained RL improvement was made possible by a strong foundation for reasoning.
Beam was pretrained in 4 weeks on 24T high-quality tokens to have innate coding capabilities.
Innovations in MoE stability and large-scale data curation & deduplication enabled us to produce a foundation that outperforms available open-source base models of the same class.
