AI may know chain-of-thought is monitored, and monitoring alone won't suffice
Original title"You could have a situation where the model understands what chain of thought is and that people are observing it. This is all in the pre...
AISummary
Dwarkesh Patel argues a model could understand that its chain of thought is being observed, since that knowledge is present in pre-training data. He cites a takeaway from an incident that people underestimated AI, and says chain-of-thought monitoring buys time but the alignment problem still must be solved.
Source: Dwarkesh Patel · x.comPublished · added here