Amanda Askell Says Aligned and Harmless Are Separate Axes in Claude Eval Incidents
AIAmanda Askell disagrees with one takeaway from Anthropic's review of Claude incidents in third-party cybersecurity evaluations. She argues models can behave in aligned ways while still causing harm, for example when given false information about their situation, because alignment and harmlessness are different axes rather than one line.
Why it matters: The author disputes the takeaway that aligned and harmless are one line, arguing they are separate axes, which sharpens how readers should interpret the evaluation incidents.



