Skip to content
Google Developers Blog·· 15d agoPickAI score62

Google reproduces Olmo 3 7B pre-training in MaxText on TPUs

Reproducing Olmo 3 7B Pre-training in MaxText: case study of large scale training on TPUs

AI summary

Google Developers reproduced Ai2's Olmo 3 7B from scratch in MaxText on Google Cloud TPUs, covering both the stage-1 pre-training run and the stage-2 mid-training anneal. The match was checked on held-out C4 loss, an 8-task accuracy suite, multi-domain perplexity, and token-level KL, not just the training loss curve. The post also describes a data-loader bug that made training loss look better than the reference while held-out metrics did not move.

Why it matters

The post documents how a faithful reproduction was verified on held-out metrics, including a data bug that training loss alone would have hidden.

Read the original developers.googleblog.com

Source: Google Developers Blog · developers.googleblog.com