OpenAI's GPT-6.1-Sol leads new Arena Alignment Index for agents
Original titleModels are getting a lot more aligned, quickly!
AISummary
The Arena Alignment Index, built from over 90K real-world agent sessions across 27 models, ranks OpenAI's GPT-6.1-Sol first with a score of 87.9, ahead of Claude-Opus-5.5 at 83.2 and Grok-4.7 at 82.7.
GPT-6.1-Sol also posted the lowest observed rates across the index's three signals: 0.89% Unauthorized Action, 1.98% False Attribution, and 2.34% Deceptive Completion.
The index's authors report that newer models consistently outperform their predecessors across all four labs, suggesting broad progress in agent safety.
Source: Boris Power · x.comPublished · added here