Alexandr Wang says no one knows how to solve AI alignment, proposes scalable oversight
Overview
Meta Chief AI Officer Alexandr Wang says nobody yet knows exactly how to solve AI alignment, which he calls one of the most open scientific questions in AI. As reported by Rohan Paul on X, Wang proposes scalable oversight, in which a separate set of AIs monitors more capable models and must improve alongside them.
Wang says Meta's Muse already uses a version of this approach, with a sentinel agent that checks the main agent's actions.
Written by AI from the articles below · updated Oct 9, 8:01 PM ET
Check the sources:
Article timeline
The articles in this story. Times are ET.
Rohan Paul@rohanpaul_aiXAlexandr Wang says nobody yet knows how to solve AI alignmentAIMeta Chief AI Officer Alexandr Wang says nobody knows exactly how to solve AI alignment, calling it one of the most open scientific questions in AI. He proposes scalable oversight, in which a separate set of AIs monitors more capable models, and says those watcher AIs must improve alongside the models they check. He adds that Meta's Muse already uses a version of this, with a sentinel agent checking the main agent's actions.

Heat trend
Not enough continuous observations to show a trend yet.