Skip to content
Read the original: OpenAI Alignment Research Blog· Micah Carroll, Tomek Korbak, Zehao Dou, Bowen Baker, Ian Kivlichan·Published· May 6, 2026PickAI score62

OpenAI finds accidental chain-of-thought grading in several RL runs but no clear monitorability loss

Investigating the consequences of accidentally grading CoT during RL

AISummary

OpenAI reports that its automated system found accidental chain-of-thought grading in RL runs for several released models, including GPT-5.4 Thinking and GPT-5.4 mini. Its analysis found no clear reduction in CoT monitorability, though the company says subtler effects cannot be ruled out. OpenAI says it still avoids grading CoTs during RL and has fixed the affected reward pathways.

AIWhy it matters

The post shows how accidental chain-of-thought grading was detected and tested, giving a concrete method for checking monitorability risks in RL training.

Read the original alignment.openai.com

Source: OpenAI Alignment Research Blog · alignment.openai.com