We frame progress on scalable oversight similar to weak-to-strong generalization: what fraction of the performance of a strong model trained on “golden” data can be recovered when training the strong model only using a weak supervision signal? This blog post explains it: https://openai.com/index/weak-to-strong-generalization/
We frame progress on scalable oversight similar to weak-to-strong generalization: what fraction of the performance of a strong model trained on “golden” data can be recovered when training the strong model only using a weak supervision signal? This blog post explains it: https://openai.com/index/weak-to-strong-generalization/