Arena reports False Attribution rates for GPT-6 Luna, Astra and Sol models
Overview
Arena (@arena) reports that GPT-6 model variants differ in how they fail on False Attribution, where an agent credits the user with a statement, request, choice, or fact that user-provided evidence contradicts.
Arena says GPT-6 Luna and Astra rarely misquoted users (15.6% and 28.6%) but often misattributed statements to them (53.1% and 48.2%). Its sibling model, GPT-6 Sol, had the highest rate of misstating the user's history, at 23.5%.
The report contains only these rates from Arena's evaluation; it does not describe the test methodology or how the rates compare with other model families.
Written by AI from the articles below · updated Oct 8, 9:01 PM ET
Check the sources:
Article timeline
Follow the coverage from different perspectives. Times are ET.
- ArenaArena reports GPT-6 model variants' false attribution rates
AIArena found that some models misquote users while others credit users with others' work in false attribution cases. GPT-6 Luna and Astra rarely misquoted users, at 15.6% and 28.6%, but often misattributed statements, at 53.1% and 48.2%. Sibling model GPT-6 Sol had the highest rate of misstating the user's history, at 23.5%.
Heat trend
Not enough continuous observations to show a trend yet.