Vals AI finds agent teams cost up to 5.1x more for barely better results
Overview
Vals AI tested GPT-6 Sol and Claude Opus 5.5 on the Vibe Code Bench as solo agents and as agent teams, at two reasoning levels.
The teams cost 1.8x to 5.1x more than single agents. Of four comparisons, only one showed a statistically significant gain: GPT-6 Sol at medium reasoning, where the team scored 7.3 points higher. At maximum reasoning, the team setup gave neither model a real advantage.
The article also cites Anthropic tests in which more agents mainly sped up results, and OpenAI researcher Noam Brown, who says multi-agent systems mostly buy speed rather than better quality.
Written by AI from the articles below · updated Oct 11, 12:37 PM ET
Check the sources:
Article timeline
The articles in this story. Times are ET.
- The DecoderNewsAI agent teams cost up to 5.1x more with barely measurable quality gains
AIVals AI tested GPT-6 Sol and Claude Opus 5.5 on the Vibe Code Bench as solo agents and teams at two reasoning levels. Teams cost 1.8x to 5.1x more, and only one of four comparisons showed a statistically significant gain, GPT-6 Sol at medium reasoning, scoring 7.3 points higher. The article also cites Anthropic tests where more agents mainly sped up results, and OpenAI researcher Noam Brown saying multi-agent systems mostly buy speed, not better quality.
Heat trend
Not enough continuous observations to show a trend yet.