Embodied AI demos are advancing rapidly, but how much do high benchmark scores actually reflect physical reality? Introducing FlagEval-Robo — an open, dual-track evaluation suite connecting simulation with real-world execution. We systematically post-trained and stress-tested 12 leading open-weight models under strictly aligned conditions. Here is what we discovered👇
Embodied AI demos are advancing rapidly, but how much do high benchmark scores actually reflect physical reality?
AISummary
Embodied AI demos are advancing rapidly, but how much do high benchmark scores actually reflect physical reality? Introducing FlagEval-Robo — an open, dual-track evaluation suite connecting simulation with real-world execution. We systematically post-trained and stress-tested 12 leading open-weight models under strictly aligned conditions. Here is what we discovered👇
Source: BAAI · x.com