Skip to content
Read the original: Thinking Machines Lab· Published Pick70/100AI score70/100

Thinking Machines Lab explains on-policy distillation for cheaper LLM post-training

Original titleOn-Policy Distillation

AISummary

Thinking Machines Lab describes on-policy distillation, which samples rollouts from a student model and has a teacher grade each token with reverse KL. The authors report that this matches Qwen3-style reasoning results at a fraction of RL's cost, with AIME'24 reaching 70% in about 150 steps from a 400k SFT checkpoint.

The method also helps recover instruction-following behavior lost during fine-tuning on internal documents.

AIWhy it matters

The post explains why on-policy distillation gives dense per-token feedback, letting a small model match RL results at much lower compute cost.

Read the original thinkingmachines.ai

Source: Thinking Machines Lab · thinkingmachines.aiPublished · added here