AI refusal as a safety measure is unreliable and risks censorship
Overview
MIT Technology Review's Arthur Holland Michel argues that AI refusal, the main safety mechanism in modern models, is an unreliable basis for safety.
He says refusal is probabilistic, can be bypassed through jailbreaks, and is hard to draw clear lines for, and that its behavior remains poorly understood.
The article also warns that refusal could be used for censorship. It cites studies showing refusal skewed toward repressive governments, and argues that governments and companies could exploit it to restrict speech.
Written by AI from the articles below · updated Oct 9, 5:48 AM ET
Check the sources:
Article timeline
Follow the coverage from different perspectives. Times are ET.
- MIT Technology Review · AIAI refusal is probabilistic and unreliable, and it raises censorship risks
AIThe article argues that AI refusal, the main safety mechanism in modern models, is unreliable and hard to draw lines for. It cites jailbreaks, classifier stacks, and studies showing refusal skewed toward repressive governments. It warns that governments and companies could use refusal to censor speech, and that refusal behavior remains poorly understood.
Heat trend
Not enough continuous observations to show a trend yet.