Kaiming He's team proposes VISTA, a visual harness for multimodal reasoning
Original title (Chinese)
何恺明团队发布多模态Harness框架
AISummary
Kaiming He's team has introduced VISTA, a visual-native harness that lets existing multimodal models observe environments, store raw frames, and revisit them during reasoning without retraining. The paper reports that with GPT-5.6 Sol, VISTA raised the GameWorld success rate from 40.0% to 63.3% and the BabyVision maze and connection task accuracy from 41.0% to 63.2%.
Source: QbitAI · qbitai.comPublished · added here