Prime Intellect tests frontier models on 153 autonomous nanoGPT research runs
AIPrime Intellect ran 153 autonomous runs on the nanoGPT optimizer speedrun across 18 frontier models, with runs lasting up to eight days on 8xH200s. The results show a large gap between models at every stage of the research process, though none of the runs produced a fundamentally new method.
Why it matters: The experiment measures how frontier models conduct autonomous research, showing large gaps between models in experiment choice, execution, and result interpretation.