Thinking Machines Lab explains on-policy distillation for cheaper LLM post-training
Original titleOn-Policy Distillation
AISummary
Thinking Machines Lab describes on-policy distillation, which samples rollouts from a student model and has a teacher grade each token with reverse KL. The authors report that this matches Qwen3-style reasoning results at a fraction of RL's cost, with AIME'24 reaching 70% in about 150 steps from a 400k SFT checkpoint.
The method also helps recover instruction-following behavior lost during fine-tuning on internal documents.
AIWhy it matters
The post explains why on-policy distillation gives dense per-token feedback, letting a small model match RL results at much lower compute cost.
Source: Thinking Machines Lab · thinkingmachines.aiPublished · added here