WildChat-based simulation predicts OpenAI production misalignment rates within roughly 3x
Original titleCan public chat data predict real-world AI misalignments?
OpenAI's alignment team found that re-generating 100,000 WildChat conversations with five recent OpenAI models predicted production failure rates across four orders of magnitude, with 95% of predictions within 1.04 orders of magnitude.
The approach was weaker for agentic misalignment categories, where errors were about 37 times larger, and it still held roughly without access to chain-of-thought reasoning, with mean multiplicative error rising from 3.6x to 4.0x.
The post tests whether public chat data can predict real production failure rates, and where that prediction breaks down for agentic behavior.
Source: OpenAI Alignment Research Blog · alignment.openai.comPublished · added here