Cognition tests OpenAI o1 models in Devin's coding agent benchmark
Original titleA review of OpenAI’s o1 and how we evaluate coding agents
AISummary
Cognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark.
The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.
AIWhy it matters
The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.
Source: Cognition Blog (Devin, Windsurf) · cognition.comPublished · added here