Amazon paper tests LLMs as predictors of which ML experiments will help
Overview
Amazon researchers report that an LLM can act as a research world model, predicting whether an experiment will improve results before it is run, according to a paper described by Rohan Paul on X. The LLM was tested on 2,653 real experiment records from 9 setups, spanning pretraining to inference.
According to the post, adding past records from the same setup raised the average ranking correlation between predictions and actual results from 0.506 to 0.774 across 5 setups. On OLMo3-100M, adding records at low reasoning effort scored 0.892, compared with 0.648 at max effort without records.
Written by AI from the articles below · updated Oct 11, 7:43 AM ET
Check the sources:
Article timeline
The articles in this story. Times are ET.
Rohan Paul@rohanpaul_aiXAmazon paper shows LLMs can predict which experiments will helpAIAmazon researchers show an LLM can act as a research world model, predicting an experiment's gain before it runs. Tested on 2,653 real experiment records from 9 setups, past records from the same setup raised average ranking correlation with actual results from 0.506 to 0.774 across 5 setups. On OLMo3-100M, adding records at low reasoning effort scored 0.892, versus 0.648 at max effort without records.

Heat trend
Not enough continuous observations to show a trend yet.