Epoch AI notes benchmark gains from on-policy self-distillation (SDPO)
AIEpoch AI reports that AI models instructed to improve on several benchmarks could already be boosted by a recent human-authored post-training method, on-policy self-distillation (SDPO). The post implies these gains were known before the models' own technique was evaluated.














