PPO's LLM-era revival and the unexpected reasons behind it
Original titlePPO had a second wave in the LLM era for reasons unanticipated by the original paper
AISummary
John Schulman says PPO gained a second wave in the LLM era for reasons not anticipated in the original paper. He points to the importance-ratio objective, which corrects biases from numeric error, asynchronous training, and forward-pass noise, and to the clipping objective, whose effect on entropy was unknown at publication, citing DAPO's arXiv paper.
Source: John Schulman · x.comPublished · added here