Dwarkesh Patel: Pretraining gains come mostly from data, not model architecture
Original titlePretraining progress seems to be coming mostly from data improvements.
AISummary
Dwarkesh Patel and a collaborator pretrained open model recipes and data corpora from 2019 to 2025 at small scales. Data improvements contributed 12.0x compute multipliers versus 3.7x for model improvements, and the two gains stack independently.
Source: Dwarkesh Patel · x.comPublished · added here