AI agent teams cost up to 5.1x more with barely measurable quality gains
Original titleAI agent teams waste massive tokens for barely measurable quality gains, research finds
AISummary
Vals AI tested GPT-6 Sol and Claude Opus 5.5 on the Vibe Code Bench as solo agents and teams at two reasoning levels.
Teams cost 1.8x to 5.1x more, and only one of four comparisons showed a statistically significant gain, GPT-6 Sol at medium reasoning, scoring 7.3 points higher.
The article also cites Anthropic tests where more agents mainly sped up results, and OpenAI researcher Noam Brown saying multi-agent systems mostly buy speed, not better quality.
Source: The Decoder · the-decoder.comPublished · added here