GPT-5.5 and Opus 4.7 Fail ARC-AGI-3 Tasks Through Flawed World Models
Original titleAnalyzing GPT-5.5 & Opus 4.7 with ARC-AGI-3
AISummary
OpenAI's GPT-5.5 scored 0.43% and Anthropic's Opus 4.7 scored 0.18% on ARC-AGI-3, a set of 135 novel environments, according to ARC Prize's replay analysis of 160 runs.
The analysis found three recurring failure modes: models perceived local action effects but failed to build global rules, mapped unfamiliar games onto known ones, and sometimes beat a level without learning the underlying mechanic.
ARC Prize is open-sourcing its analysis package.
Source: ARC Prize · arcprize.org