Skip to content
Read the original: OpenAI Alignment Research Blog· Published Pick60/100AI score60/100

WildChat-based simulation predicts OpenAI production misalignment rates within roughly 3x

Original titleCan public chat data predict real-world AI misalignments?

AISummary

OpenAI's alignment team found that re-generating 100,000 WildChat conversations with five recent OpenAI models predicted production failure rates across four orders of magnitude, with 95% of predictions within 1.04 orders of magnitude.

The approach was weaker for agentic misalignment categories, where errors were about 37 times larger, and it still held roughly without access to chain-of-thought reasoning, with mean multiplicative error rising from 3.6x to 4.0x.

AIWhy it matters

The post tests whether public chat data can predict real production failure rates, and where that prediction breaks down for agentic behavior.

Read the original alignment.openai.com

Source: OpenAI Alignment Research Blog · alignment.openai.comPublished · added here