Read the original: OpenAI Alignment Research Blog· Micah Carroll, Tomek Korbak, Zehao Dou, Bowen Baker, Ian Kivlichan·Published· May 6, 2026PickAI score62
OpenAI finds accidental chain-of-thought grading in several RL runs but no clear monitorability loss
Investigating the consequences of accidentally grading CoT during RL
AISummary
OpenAI reports that its automated system found accidental chain-of-thought grading in RL runs for several released models, including GPT-5.4 Thinking and GPT-5.4 mini. Its analysis found no clear reduction in CoT monitorability, though the company says subtler effects cannot be ruled out. OpenAI says it still avoids grading CoTs during RL and has fixed the affected reward pathways.
AIWhy it matters
The post shows how accidental chain-of-thought grading was detected and tested, giving a concrete method for checking monitorability risks in RL training.
Source: OpenAI Alignment Research Blog · alignment.openai.com