Amanda Askell Says Aligned and Harmless Are Separate Axes in Claude Eval Incidents
Original titleI don't agree with this part. I think the takeaway should be that models (like humans) can behave in aligned ways while still causing har...
AISummary
Amanda Askell disagrees with one takeaway from Anthropic's review of Claude incidents in third-party cybersecurity evaluations. She argues models can behave in aligned ways while still causing harm, for example when given false information about their situation, because alignment and harmlessness are different axes rather than one line.
Source: Amanda Askell · x.com