Skip to content

Formats · Latest news

Expert opinion

Views from founders, researchers, investors, and other consequential voices in AI.

63 top picks all-time · 23 in the past 30 days · chosen from 1,646 items collected all-time

Latest pick

Top picks archive · Page 4

Top picks 61–63 of 63

Nov 13, 2025

Nov 13, 2025Thu
  1. Cognition Blog (Devin, Windsurf)OfficialAI score65

    Cognition's Devin review says it excels at scoped junior-level engineering work

    AICognition's 2025 performance review says Devin works best on clear, verifiable tasks such as migrations, vulnerability fixes, and unit tests. The company reports a 67% PR merge rate, up from 34% last year, and cites a bank that cut migration time per file from 30-40 hours to 3-4 hours. It also says Devin struggles with ambiguous requirements, mid-task scope changes, and soft-skill work that still needs human engineers.

    Why it matters: The report pairs concrete migration, vulnerability, and test-coverage figures with named weaknesses, letting engineering leaders judge where an agent fits in their own workflow.

Jun 11, 2025

Jun 11, 2025Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition argues multi-agent architectures are fragile and proposes context-sharing principles

    AICognition argues that parallel multi-agent architectures are fragile because subagents act on conflicting, unshared assumptions. It proposes two principles for reliable agents: share context and full agent traces, and treat actions as carrying implicit decisions. The post recommends simpler single-threaded designs for most cases and notes that context compression and fine-tuned models can extend long-running tasks.

    Why it matters: The post explains concrete failure modes of parallel multi-agent setups and offers two context-sharing principles, useful for anyone designing long-running agent systems.

Sep 11, 2024

Sep 11, 2024Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score60

    Cognition tests OpenAI o1 models in Devin's coding agent benchmark

    AICognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.

    Why it matters: The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.