Was really interesting to hear John, Beren, and Charlie speculate about why Sonnet 5 and Opus 5 feel like worse models than GLM 5.3 (despite the fact that Anthropic can do raw logit distillation from…
Original titleWas really interesting to hear John, Beren, and Charlie speculate about why Sonnet 5 and Opus 5 feel like worse models than GLM 5.3
AISummary
…Fable, and can also train Sonnet/Opus on the environments from which Fable was trained). Led to some interesting thoughts about value of distillation, what it takes to do distillation effectively, and what kinds of model behaviors are hard to extract from distillation.
Source: Dwarkesh Patel · x.com