Skip to content

Companies & models · Latest news

OpenAI / ChatGPT

Follow GPT models, ChatGPT and Sora products, company strategy, and personnel at OpenAI.

101 top picks all-time · 59 in the past 30 days · chosen from 739 items collected all-time

Latest pick

Top picks archive · Page 6

Top picks 101–101 of 101

Sep 11, 2024

Sep 11, 2024Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score60

    Cognition tests OpenAI o1 models in Devin's coding agent benchmark

    AICognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.

    Why it matters: The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.