We frame progress on scalable oversight similar to weak-to-strong generalization: what fraction of the performance of a strong model trained on “golden” data can be recovered when training the strong…
Original titleWe frame progress on scalable oversight similar to weak-to-strong generalization: what fraction of the performance of a strong model trai...
AISummary
…model only using a weak supervision signal? This blog post explains it:
Source: Jan Leike · x.com