However, most alignment research is not very crisp and requires research taste when evaluating.
Original titleHowever, most alignment research is not very crisp and requires research taste when evaluating.
AISummary
This is why we chose to point the AAR at this scalable oversight problem! Progress would let AARs work on fuzzier alignment problems, where humans can only provide weak supervision.
Source: Jan Leike · x.com