Skip to content
View original post on X: Rohan PaulX· 40/100AI score40/100

Alexandr Wang says nobody yet knows how to solve AI alignment

AISummary

Meta Chief AI Officer Alexandr Wang says nobody knows exactly how to solve AI alignment, calling it one of the most open scientific questions in AI. He proposes scalable oversight, in which a separate set of AIs monitors more capable models, and says those watcher AIs must improve alongside the models they check.

He adds that Meta's Muse already uses a version of this, with a sentinel agent checking the main agent's actions.

Post on XView on X
Rohan PaulVerified on X
@rohanpaul_ai

Alexandr Wang ( @alexandr_wang, Chief AI Officer at Meta): nobody knows how to solve alignment yet

"This is, I think, one of the most open questions scientifically in AI. I think nobody knows exactly the way to solve this problem, but there’s a few ideas."

His solution is "scalable oversight":

“as the AIs get smarter, we use a different set of AIs to observe what they’re doing and keep them in check.”

The watcher AIs have to improve along with the models they watch, so labs would need to build “smarter and smarter policing agents” too.

Meta’s Muse already uses a version of this, with a separate sentinel agent checking what the main agent does.

----
Full video on "Cleo Abram" YouTube channel, (link in comment)

Source: Rohan Paul · x.comPublished · added here