Skip to content

Areas · Latest news

AI agents

Models that plan, use tools, and complete multistep tasks, from Claude Code and Manus to agent frameworks and evaluations.

267 top picks all-time · 123 in the past 30 days · chosen from 1,899 items collected all-time

Latest pick

Top picks archive · Page 14

Top picks 261–267 of 267

Jun 11, 2025

Jun 11, 2025Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition argues multi-agent architectures are fragile and proposes context-sharing principles

    AICognition argues that parallel multi-agent architectures are fragile because subagents act on conflicting, unshared assumptions. It proposes two principles for reliable agents: share context and full agent traces, and treat actions as carrying implicit decisions. The post recommends simpler single-threaded designs for most cases and notes that context compression and fine-tuned models can extend long-running tasks.

    Why it matters: The post explains concrete failure modes of parallel multi-agent setups and offers two context-sharing principles, useful for anyone designing long-running agent systems.

May 14, 2025

May 14, 2025Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Devin 2.1 adds confidence ratings and built-in codebase intelligence

    AICognition has released Devin 2.1, which reports its confidence in completing tasks using green, yellow, and red ratings. The company says green scores led to twice the likelihood of a merged PR compared with red, and Devin now also answers codebase questions and scores Linear and Jira issues.

    Why it matters: The post explains how Devin now shows confidence scores and asks clarifying questions, which changes how teams can decide which tasks to hand over.

Apr 2, 2025

Apr 2, 2025Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score75

    Cognition launches Devin 2.0 with agent-native IDE and new planning tools

    AICognition has released Devin 2.0, a new agent-native IDE experience with a flexible plan starting at $20. The update lets users run multiple parallel Devins, each with its own cloud-based IDE, and adds Interactive Planning, Devin Search, and Devin Wiki.

    Why it matters: The release adds planning, codebase search, and auto-generated wikis to Devin, showing how an agent can prepare work before executing it.

Dec 9, 2024

Dec 9, 2024Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score67

    Cognition makes Devin generally available to engineering teams from $500 a month

    AICognition is making Devin generally available to engineering teams starting at $500 a month, with no seat limits and access to its Slack integration, IDE extension, and API. The post recommends starting with small frontend bugs, first-draft PRs for backlog tasks, and targeted refactors, and shares open-source PR sessions where Devin resolved issues for projects including Anthropic MCP, Zod, and nanoGPT.

    Why it matters: The post shows concrete open-source PR examples and the tasks where Devin works best, helping teams judge where an autonomous coding agent fits their workflow.

Sep 11, 2024

Sep 11, 2024Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score60

    Cognition tests OpenAI o1 models in Devin's coding agent benchmark

    AICognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.

    Why it matters: The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.

Mar 14, 2024

Mar 14, 2024Thu
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition reports Devin resolves 13.86% of SWE-bench issues end to end

    AICognition reports that its agent Devin resolved 79 of 570 sampled SWE-bench issues, a 13.86% success rate, without being given the files to edit. The report says this exceeds the best previous unassisted baseline of 1.96% and the best assisted result of 4.80%. It also describes the adapted evaluation setup, a 45-minute runtime limit, and cases where Devin failed on multi-file edits.

    Why it matters: The report explains how SWE-bench was adapted for end-to-end agent evaluation, with failure cases that clarify where the 13.86% result comes from and its limits.

Mar 11, 2024

Mar 11, 2024Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score88

    Cognition introduces Devin, an AI agent that works on software engineering tasks

    AICognition introduces Devin as an AI software engineer that can plan and execute complex engineering tasks with a shell, code editor, and browser. On SWE-bench, Devin resolved 13.86% of issues end-to-end, versus a previous state-of-the-art of 1.96%, on a random 25% subset of the dataset. Devin is in early access, with access available through a waitlist.

    Why it matters: The post pairs Devin's end-to-end task demos with SWE-bench results against prior models, letting readers weigh the claimed capability against the evaluation setup.