Thinking Machines explores on-policy distillation for training small models
Original titleCombining the benefits of RL and SFT with on-policy distillation, a promising approach for training small models for domain performance a...
AISummary
Thinking Machines published a post on on-policy distillation, a training approach combining the error-correcting relevance of RL with the reward density of SFT. The quoted post reports that in math reasoning and an internal chat assistant, on-policy distillation can outperform other approaches at a fraction of the cost.
Source: Mira Murati · x.comPublished · added here