Kaiming He's team proposes VISTA, a visual harness for multimodal reasoning in interactive worlds
Overview
Kaiming He's team, at MIT, has introduced VISTA, a visual-native harness that lets existing multimodal models observe environments, store raw frames, and revisit them during reasoning without retraining.
The harness works with already-trained models and does not require training them from scratch.
The paper reports that with GPT-5.6 Sol, VISTA raised the GameWorld success rate from 40.0% to 63.3%. On the BabyVision maze and connection task, accuracy rose from 41.0% to 63.2%. These results come from the paper as reported by QbitAI, and no independent evaluation is cited.
Written by AI from the articles below · updated Oct 11, 7:38 AM ET
Check the sources:
Article timeline
The articles in this story. Times are ET.
- QbitAINewsKaiming He's team proposes VISTA, a visual harness for multimodal reasoning
AIKaiming He's team has introduced VISTA, a visual-native harness that lets existing multimodal models observe environments, store raw frames, and revisit them during reasoning without retraining. The paper reports that with GPT-5.6 Sol, VISTA raised the GameWorld success rate from 40.0% to 63.3% and the BabyVision maze and connection task accuracy from 41.0% to 63.2%.
Heat trend
Not enough continuous observations to show a trend yet.