AI refusal is probabilistic and unreliable, and it raises censorship risks
AIThe article argues that AI refusal, the main safety mechanism in modern models, is unreliable and hard to draw lines for. It cites jailbreaks, classifier stacks, and studies showing refusal skewed toward repressive governments. It warns that governments and companies could use refusal to censor speech, and that refusal behavior remains poorly understood.