Claude's scalable oversight methods generalize well to math but poorly to code
AIAnthropic's Claude developed scalable oversight methods on chat reward modeling datasets and evaluated them on math and code datasets. The best methods performed strongly on math but gave mixed results on code, suggesting the methods were overfit to the data and models used.






