Skip to contentSkip to stories

Updated

#Agent

Nov 3, 2025

Nov 3, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score47

    Windsurf Codemaps Adds AI-Annotated Code Maps to Help Engineers Understand Code

    AIWindsurf has launched Codemaps, AI-annotated structured maps of a codebase powered by SWE-1.5 and Claude Sonnet 4.5, which users can generate from a task prompt using a Fast (SWE-1.5) or Smart (Sonnet 4.5) model. Codemaps links grouped code sections to exact lines and can be referenced in Cascade with @{codemap} to give agents more specific context.

Oct 30, 2025

Oct 30, 2025Thu
  1. Chip HuyenAI score27

    Chip Huyen's AI product lessons: UX, data, and team structure matter most

    AIChip Huyen argues that many AI product failures stem from user experience, data quality, and organizational structure rather than the AI itself. She cites a chatbot whose traction improved after adding pre-populated questions and a voice option for users whose hands were busy, and a lead scoring model that was broken because marketing wasn't asking the right questions. She also notes that senior engineers gain the most from AI coding while resisting it more, and recommends building small tools for daily frustrations to solve the "idea crisis."

Oct 28, 2025

Oct 28, 2025Tue
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition releases SWE-1.5, a coding agent model served at up to 950 tok/s

    AICognition has released SWE-1.5, a model optimized for software engineering that it says reaches near-frontier coding performance while running at up to 950 tok/s with Cerebras inference. The company reports it is 6x faster than Haiku 4.5 and 13x faster than Sonnet 4.5, and it is available now in Windsurf. The post's SWE-Bench Pro chart places SWE-1.5 at 40.08%, behind Sonnet 4.5 at 43.60%, and it notes that the model was trained with reinforcement learning on the Cascade agent harness.

    Why it matters: The post pairs a benchmark chart with a 950 tok/s speed claim and describes how harness, RL environments, and inference were co-designed, useful context for judging the speed-versus-quality tradeoff.

Oct 27, 2025

Oct 27, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score36

    Devin Automates .NET Framework to .NET Core Migration in Weeks, Not Months

    AICognition says its autonomous coding agent Devin can complete a .NET Framework to .NET Core migration in as little as two weeks, using a Strangler Fig approach adapted from Jimmy Bogard's guide. The post says Devin handles planning via Ask Devin and DeepWiki, dependency sharing, controller and view conversion, and session state adaptation through a remote app.

Oct 26, 2025

Oct 26, 2025Sun
  1. Factory NewsAI score36

    AWS and Factory Announce Partnership, Factory Available on AWS Marketplace

    AIFactory has announced a partnership with Amazon Web Services and made its Droids agent platform available on the AWS Marketplace. Enterprise teams can use existing AWS Enterprise Discount Program commitments to buy Factory, with Droids accessible from CLI, Terminal UI, Web, Slack, Linear, and an IDE overlay. The source cites 31× faster feature development, 96.1%+ reduction in migration times, and 95.8% reduction in incident resolution times.

Oct 15, 2025

Oct 15, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score73

    Cognition releases SWE-grep models for fast parallel code context retrieval

    AICognition introduces SWE-grep and SWE-grep-mini, fast agentic models trained with reinforcement learning for multi-turn context retrieval in coding tasks. The company says they match frontier coding models at retrieval while taking an order of magnitude less time, and they power the Fast Context subagent in Windsurf. The models issue up to 8 parallel tool calls per turn within 4 turns, and Cerebras serves SWE-grep-mini at over 2,800 tokens per second and SWE-grep at over 650 tokens per second.

    Why it matters: The post explains the speed-intelligence tradeoff in agentic code search, showing how parallel tool calls and RL training change the cost of retrieving context for coding agents.

Sep 28, 2025

Sep 28, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition rebuilds Devin around Claude Sonnet 4.5 for 2x speed

    AICognition rebuilt its Devin coding agent for Claude Sonnet 4.5, reporting 2x faster performance and 12% better results on its Junior Developer Evals, now available in Agent Preview. The team found the model is aware of its context window, which led to premature wrap-up behavior that they countered with repeated prompts and a 200k usage cap within a 1M token beta.

    Why it matters: The post explains which agent behaviors changed under Sonnet 4.5, such as context-window awareness and note-taking, that forced a rebuild rather than a simple model swap.

Sep 7, 2025

Sep 7, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score53

    Cognition raises over $400M at $10.2B valuation after Windsurf acquisition

    AICognition, maker of the AI software engineer Devin, raised over $400M at a $10.2B post-money valuation led by Founders Fund. The company says its acquisition of Windsurf more than doubled its ARR, with combined enterprise ARR up over 30% in the seven weeks after the deal. It also reports Devin ARR grew from $1M in September 2024 to $73M in June 2025, with total net burn under $20M.

Sep 3, 2025

Sep 3, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score38

    Eight Sleep Uses Devin AI as Data Analyst to Clear Ad-Hoc Requests

    AIEight Sleep integrated Cognition's Devin into its data workflows, letting staff tag Devin in Slack to query Snowflake, dbt, and Looker and check Amplitude. The company says it is now shipping 3x as many data features and investigations each week, with its ad-hoc data request queue near zero. Devin was used to trace a suspicious revenue spike to a better-than-expected email campaign.

Aug 27, 2025

Aug 27, 2025Wed

Aug 4, 2025

Aug 4, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score22

    Devin Can Automate Migrating Jenkins Pipelines to GitHub Actions at Scale

    AICognition says generative AI agents such as Devin can read internal docs and convert Jenkins pipelines to GitHub Actions syntax, replacing custom plugins with Actions or APIs and validating the results. The company claims enterprises can cut multi-year migration efforts to a few months, with Devin running inside the customer's secure environment so code and secrets stay internal.

Jul 21, 2025

Jul 21, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score50

    Devin Adds MCP Support and a Marketplace for Connecting External Servers

    AICognition's Devin is now compatible with the Model Context Protocol (MCP), letting users connect favorite MCP servers through a new MCP Marketplace found in Settings. The source cites example uses including querying Datadog and Sentry logs, creating Notion docs, Google Docs, and Linear tickets, and interacting with Figma, Airtable, Stripe, and Hubspot.

Jun 26, 2025

Jun 26, 2025Thu

Jun 11, 2025

Jun 11, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition argues multi-agent architectures are fragile and proposes context-sharing principles

    AICognition argues that parallel multi-agent architectures are fragile because subagents act on conflicting, unshared assumptions. It proposes two principles for reliable agents: share context and full agent traces, and treat actions as carrying implicit decisions. The post recommends simpler single-threaded designs for most cases and notes that context compression and fine-tuned models can extend long-running tasks.

    Why it matters: The post explains concrete failure modes of parallel multi-agent setups and offers two context-sharing principles, useful for anyone designing long-running agent systems.

May 14, 2025

May 14, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score62

    Devin 2.1 adds confidence ratings and built-in codebase intelligence

    AICognition has released Devin 2.1, which reports its confidence in completing tasks using green, yellow, and red ratings. The company says green scores led to twice the likelihood of a merged PR compared with red, and Devin now also answers codebase questions and scores Linear and Jira issues.

    Why it matters: The post explains how Devin now shows confidence scores and asks clarifying questions, which changes how teams can decide which tasks to hand over.

Apr 2, 2025

Apr 2, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score75

    Cognition launches Devin 2.0 with agent-native IDE and new planning tools

    AICognition has released Devin 2.0, a new agent-native IDE experience with a flexible plan starting at $20. The update lets users run multiple parallel Devins, each with its own cloud-based IDE, and adds Interactive Planning, Devin Search, and Devin Wiki.

    Why it matters: The release adds planning, codebase search, and auto-generated wikis to Devin, showing how an agent can prepare work before executing it.

Feb 25, 2025

Feb 25, 2025Tue
  1. Cognition Blog (Devin, Windsurf)AI score40

    Devin Gets Batch Edits, GitLab Support, and Sonnet 3.7 in February Update

    AICognition's February 2025 Devin update adds parallel batch edits, beta GitLab support, and Sonnet 3.7, which Cognition says is the best model it has tested for debugging, codebase search, and agentic planning. Devin is about 2x faster than in October 2024, taking about 7.8 minutes on average to complete junior developer tasks in Cognition's internal evaluations. Other changes include copy-paste in Devin's browser and proactive feedback on suboptimal prompts.

Feb 12, 2025

Feb 12, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score36

    Linktree Uses Devin to Add Social Platforms and Ship About 100 Merged PRs

    AILinktree has used Devin, Cognition's AI software engineer, to merge roughly 100 pull requests in a month, mostly fixing customer-reported bugs and implementing small features. The engineering team also used Devin to add support for new social media platforms, launching five Devins, one per repo and PR, and later used the Devin API with a Playbook script to spawn multiple Devins for multi-repo features. The team says Devin works best on tasks an engineer could finish in a couple of hours.

Jan 20, 2025

Jan 20, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score46

    Devin 101: Automatic PR Reviews with the Devin API

    AICognition's Devin can be triggered through its External API by GitHub Actions to automatically review pull requests, typically within five to ten minutes. The setup involves adding a workflow file, storing a DEVIN_API_KEY secret, and customizing the review prompt to match team conventions. Cognition recommends treating Devin as an extra reviewer rather than a replacement for human oversight, since it does not catch every bug.

Jan 15, 2025

Jan 15, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score40

    Devin's January 2025 Update Adds Repo Context, Enterprise Accounts, and Usage Billing

    AICognition's January 2025 Devin update improves its ability to find relevant files and reuse existing code in repositories, with changes rolling out to all users. It also adds enterprise accounts for centralized management of multiple organizations, audio message support in Slack, and pay-as-you-go billing after monthly ACU capacity is used, starting January 9.

Jan 13, 2025

Jan 13, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score38

    Crossmint Uses Devin to Scale Open-Source Development of GOAT SDK

    AICrossmint said Devin became its top contributor to the open-source GOAT SDK during an initial trial, merging 8 pull requests versus 4 for the next contributor. Examples included a DEXScreener plugin built from a documentation URL and a harder Sui blockchain integration that needed three rounds of feedback and about an hour of human involvement. The company said its results depended on proper training, clear task context, and planned validation, not on treating Devin as superhuman.

Dec 22, 2024

Dec 22, 2024Sun
  1. Cognition Blog (Devin, Windsurf)AI score57

    Devin becomes generally available with faster runs and new customization options

    AICognition has made Devin generally available to all engineering teams, with subscriptions starting at $500 per month. Over the past two weeks, Devin was made about 10% faster and about 10% more cost-efficient, especially for tasks requiring many code edits. The update also adds fixes for stuck or hanging sessions, more options to customize filters and Slack notifications, and larger machine settings for disk, RAM, and CPU.

Dec 11, 2024

Dec 11, 2024Wed
  1. Cognition Blog (Devin, Windsurf)AI score38

    Devin Open Source Initiative Gives Maintainers 500 Free ACUs for Repo Work

    AICognition is launching the Devin Open Source Initiative, giving selected open source maintainers 500 free ACUs on a Devin Teams plan as part of Devin's general availability launch. The post shows Devin contributing pull requests to projects including Anthropic's MCP Inspector, Dagger, and nanoGPT, with maintainers still reviewing the results. Devin's GitHub integration forwards PR comments and CI checks to help refine changes, though the company warns a human should still verify final quality.

Dec 9, 2024

Dec 9, 2024Mon
  1. Cognition Blog (Devin, Windsurf)AI score67

    Cognition makes Devin generally available to engineering teams from $500 a month

    AICognition is making Devin generally available to engineering teams starting at $500 a month, with no seat limits and access to its Slack integration, IDE extension, and API. The post recommends starting with small frontend bugs, first-draft PRs for backlog tasks, and targeted refactors, and shares open-source PR sessions where Devin resolved issues for projects including Anthropic MCP, Zod, and nanoGPT.

    Why it matters: The post shows concrete open-source PR examples and the tasks where Devin works best, helping teams judge where an autonomous coding agent fits their workflow.

Dec 2, 2024

Dec 2, 2024Mon
  1. Cognition Blog (Devin, Windsurf)AI score46

    Devin's December 2024 update adds Slack collaboration, PR tools, and a REST API

    AICognition's December 2024 Devin update lets users tag Devin in Slack threads and ask it to create PRs, with Devin automatically responding to PR comments and lint failures. The update adds Repo Knowledge that Devin generates by scanning repositories, an Agency setting that makes Devin propose plans before executing complex tasks, and a REST API for structured input and output. Devin also gained faster session startup and enterprise options such as Okta single sign-on.

Sep 11, 2024

Sep 11, 2024Wed
  1. Cognition Blog (Devin, Windsurf)AI score60

    Cognition tests OpenAI o1 models in Devin's coding agent benchmark

    AICognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.

    Why it matters: The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.

Sep 4, 2024

Sep 4, 2024Wed
  1. Cognition Blog (Devin, Windsurf)AI score45

    Devin Adds MultiDevin, Auto PR Replies, and Rollback in September 2024 Update

    AICognition's Devin gained a September 2024 update with MultiDevin, which lets a manager Devin delegate work to up to 10 worker Devins, currently available on the Enterprise plan. Devin now automatically responds to comments on its pull requests, suggests Knowledge additions, and can restore earlier checkpoints, and the company reports up to an 80% reduction in time for common tasks.

Jul 19, 2024

Jul 19, 2024Fri
  1. Cognition Blog (Devin, Windsurf)AI score47

    Devin Automates Recovery of Windows Machines From CrowdStrike Outage Failure

    AICognition tested whether its Devin AI agent could recover Windows machines stuck in the Blue Screen of Death after the CrowdStrike outage. Following a playbook of eight steps, Devin mounted the drive, deleted the faulty CrowdStrike files, and debugged a failed volume detach. The blog says the machine was confirmed bootable, and that playbooks are most useful when the same fix must be repeated across many machines.

Jun 4, 2024

Jun 4, 2024Tue
  1. Cognition Blog (Devin, Windsurf)AI score55

    Devin June 2024 update adds playbooks, snapshots, and Slack, Github, and Linear triggers

    AIDevin's June 2024 update lets users directly operate its VSCode editor, terminal, and browser, and adds Playbooks for recurring engineering tasks. It also adds Knowledge sharing, Machine Snapshots that persist dev environments, event-driven triggers from Slack, GitHub, and Linear, a Secrets manager, and tools for verifying Devin's work. Access remains limited to a waitlist with weekly invite releases.

Mar 14, 2024

Mar 14, 2024Thu
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition reports Devin resolves 13.86% of SWE-bench issues end to end

    AICognition reports that its agent Devin resolved 79 of 570 sampled SWE-bench issues, a 13.86% success rate, without being given the files to edit. The report says this exceeds the best previous unassisted baseline of 1.96% and the best assisted result of 4.80%. It also describes the adapted evaluation setup, a 45-minute runtime limit, and cases where Devin failed on multi-file edits.

    Why it matters: The report explains how SWE-bench was adapted for end-to-end agent evaluation, with failure cases that clarify where the 13.86% result comes from and its limits.

Mar 11, 2024

Mar 11, 2024Mon
  1. Cognition Blog (Devin, Windsurf)AI score88

    Cognition introduces Devin, an AI agent that works on software engineering tasks

    AICognition introduces Devin as an AI software engineer that can plan and execute complex engineering tasks with a shell, code editor, and browser. On SWE-bench, Devin resolved 13.86% of issues end-to-end, versus a previous state-of-the-art of 1.96%, on a random 25% subset of the dataset. Devin is in early access, with access available through a waitlist.

    Why it matters: The post pairs Devin's end-to-end task demos with SWE-bench results against prior models, letting readers weigh the claimed capability against the evaluation setup.