Thinking Machines explores on-policy distillation for training small models
AIThinking Machines published a post on on-policy distillation, a training approach combining the error-correcting relevance of RL with the reward density of SFT. The quoted post reports that in math reasoning and an internal chat assistant, on-policy distillation can outperform other approaches at a fraction of the cost.