@SpaceXAI Grok 4.7's reasoning-token usage correlates with its ARC-AGI-2 public score: low averaged 10k tokens per test-pair attempt and scored 25%, versus 86k to 120k tokens and 57.5% to 60% at…
Original title@SpaceXAI Grok 4.7's reasoning-token usage correlates with its ARC-AGI-2 public score: low averaged 10k tokens per test-pair attempt and ...
AISummary
…medium through xhigh. Low's lower token usage may help explain its lower score.
Source: ARC Prize · x.com