Jan Leike says alignment research needs taste, so AARs target scalable oversight
AIMost alignment research is not crisp and requires research taste to evaluate, according to Jan Leike. He explains that Anthropic chose to point the AAR at a scalable oversight problem because progress there could let AARs tackle fuzzier alignment problems where humans can only provide weak supervision.