Artificial Analysis: output token volume does not track benchmark scores across models
Overview
Artificial Analysis reports that generating more output tokens per task does not reliably produce higher benchmark scores.
In its comparison, GPT-6 Astra (max) scored 8.6% on about 81k output tokens per task, less than half the roughly 180k used by Grok 4.7 (xhigh). Three Claude models produced the most output, about 202k to 562k tokens per task, yet scored between 2.8% and 6.4%.
The finding is a single observation from Artificial Analysis's account on X, based on the token counts and scores it reports for these five models; it does not establish a general relationship between verbosity and accuracy.
AIWritten by AI from the articles below · overview updated Oct 8, 8:51 PM ET
Check the sources:
Article timeline
Follow the coverage from different perspectives. Times are ET.
- Artificial AnalysisGenerating more output tokens doesn’t necessarily translate to a higher score. GPT-6 Astra (max) scores 8.6% on ~81k output tokens per task, under half the ~180k of Grok 4.7 (xhigh). Three Claude models generated the most output tokens (~202k to ~562k per task) and score 2.8% to 6.4%.
Generating more output tokens doesn’t necessarily translate to a higher score. GPT-6 Astra (max) scores 8.6% on ~81k output tokens per task, under half the ~180k of Grok 4.7 (xhigh). Three Claude models generated the most output tokens (~202k to ~562k per task) and score 2.8% to 6.4%.
Heat trend
Not enough continuous observations to show a trend yet.