Skip to content

Areas

AI coding Latest news

Coding assistants, vibe coding, code model evaluations, and changes to software development workflows.

114 picksPast 30 days: 35 itemsTotal: 719 items

Latest pick

Top picks archive · Page 6

Feb 4

Feb 4WedItems 101–114
  1. Anthropic EngineeringAI score75

    Anthropic details how parallel Claude agents built a 100,000-line C compiler

    Nicholas Carlini of Anthropic's Safeguards team describes an agent-team setup where 16 Claude instances worked in parallel on a shared codebase without human intervention to write a Rust-based C compiler. Over nearly 2,000 Claude Code sessions costing about $20,000 in API fees, the team produced a 100,000-line compiler that can build Linux 6.9 on x86, ARM, and RISC-V. The post focuses on harness design, including high-quality tests, lock files for task claiming, GCC as a reference oracle for the kernel, and the limits the project reached.

    AIWhy it matters: The post shows concrete harness design choices for long-running agent teams, including test design, locking, and parallel work division, that readers can adapt to their own autonomous projects.

Dec 20, 2025

Dec 20, 2025Sat
  1. MiniMax · new models on Hugging FaceAI score74

    MiniMax-M2.1 open-sources weights for coding and agent tasks

    MiniMax has released MiniMax-M2.1 model weights on Hugging Face, with API access on the MiniMax Open Platform and the MiniMax Agent product. The company reports gains over M2 on coding and agent benchmarks such as SWE-bench Verified (74.0) and VIBE average (88.6), and says it outperforms Claude Sonnet 4.5 on multilingual scenarios.

    AIWhy it matters: The release pairs open weights with a broad benchmark table against Claude and GPT models, letting readers compare coding and agent claims directly.

Dec 16, 2025

Dec 16, 2025Tue
  1. Xiaomi MiMoAI score78

    Xiaomi releases open-source MiMo-V2-Flash MoE model for reasoning and coding

    Xiaomi released and open-sourced MiMo-V2-Flash, a Mixture-of-Experts model with 309B total and 15B active parameters, under the MIT license. The company reports 73.4% on SWE-Bench Verified, the top score among open-source models, and inference at 150 tokens per second for $0.1 per million input tokens and $0.3 per million output tokens. It supports a hybrid thinking mode and a 256k context window.

    AIWhy it matters: The post gives architecture, speculative decoding speedup, and pricing figures, which help readers judge how the efficiency claims are achieved and what they cost.

Dec 11, 2025

Dec 11, 2025Thu
  1. Nick TurleyAI score78

    OpenAI introduces GPT-5.2 in ChatGPT for professional work

    OpenAI is introducing GPT-5.2 in ChatGPT, describing it as its most advanced model series for professional work. GPT-5.2 Thinking is positioned for tasks such as building spreadsheets and presentations, writing and reviewing production code, and analyzing long documents. The post says it beats or ties industry professionals on well-specified knowledge work tasks spanning 44 occupations 70.9% of the time on GDPval, and GPT-5.2 Instant, Thinking, and Pro begin rolling out to all tiers, starting with paid plans.

    AIWhy it matters: The post links the model's professional-work focus to GDPval results across 44 occupations, showing how the claimed capability was measured.

Nov 13, 2025

Nov 13, 2025Thu
  1. Cognition Blog (Devin, Windsurf)AI score65

    Cognition's Devin review says it excels at scoped junior-level engineering work

    Cognition's 2025 performance review says Devin works best on clear, verifiable tasks such as migrations, vulnerability fixes, and unit tests. The company reports a 67% PR merge rate, up from 34% last year, and cites a bank that cut migration time per file from 30-40 hours to 3-4 hours. It also says Devin struggles with ambiguous requirements, mid-task scope changes, and soft-skill work that still needs human engineers.

    AIWhy it matters: The report pairs concrete migration, vulnerability, and test-coverage figures with named weaknesses, letting engineering leaders judge where an agent fits in their own workflow.

Oct 28, 2025

Oct 28, 2025Tue
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition releases SWE-1.5, a coding agent model served at up to 950 tok/s

    Cognition has released SWE-1.5, a model optimized for software engineering that it says reaches near-frontier coding performance while running at up to 950 tok/s with Cerebras inference. The company reports it is 6x faster than Haiku 4.5 and 13x faster than Sonnet 4.5, and it is available now in Windsurf. The post's SWE-Bench Pro chart places SWE-1.5 at 40.08%, behind Sonnet 4.5 at 43.60%, and it notes that the model was trained with reinforcement learning on the Cascade agent harness.

    AIWhy it matters: The post pairs a benchmark chart with a 950 tok/s speed claim and describes how harness, RL environments, and inference were co-designed, useful context for judging the speed-versus-quality tradeoff.

Oct 15, 2025

Oct 15, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score73

    Cognition releases SWE-grep models for fast parallel code context retrieval

    Cognition introduces SWE-grep and SWE-grep-mini, fast agentic models trained with reinforcement learning for multi-turn context retrieval in coding tasks. The company says they match frontier coding models at retrieval while taking an order of magnitude less time, and they power the Fast Context subagent in Windsurf. The models issue up to 8 parallel tool calls per turn within 4 turns, and Cerebras serves SWE-grep-mini at over 2,800 tokens per second and SWE-grep at over 650 tokens per second.

    AIWhy it matters: The post explains the speed-intelligence tradeoff in agentic code search, showing how parallel tool calls and RL training change the cost of retrieving context for coding agents.

Sep 28, 2025

Sep 28, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition rebuilds Devin around Claude Sonnet 4.5 for 2x speed

    Cognition rebuilt its Devin coding agent for Claude Sonnet 4.5, reporting 2x faster performance and 12% better results on its Junior Developer Evals, now available in Agent Preview. The team found the model is aware of its context window, which led to premature wrap-up behavior that they countered with repeated prompts and a 200k usage cap within a 1M token beta.

    AIWhy it matters: The post explains which agent behaviors changed under Sonnet 4.5, such as context-window awareness and note-taking, that forced a rebuild rather than a simple model swap.

May 14, 2025

May 14, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score62

    Devin 2.1 adds confidence ratings and built-in codebase intelligence

    Cognition has released Devin 2.1, which reports its confidence in completing tasks using green, yellow, and red ratings. The company says green scores led to twice the likelihood of a merged PR compared with red, and Devin now also answers codebase questions and scores Linear and Jira issues.

    AIWhy it matters: The post explains how Devin now shows confidence scores and asks clarifying questions, which changes how teams can decide which tasks to hand over.

Apr 2, 2025

Apr 2, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score75

    Cognition launches Devin 2.0 with agent-native IDE and new planning tools

    Cognition has released Devin 2.0, a new agent-native IDE experience with a flexible plan starting at $20. The update lets users run multiple parallel Devins, each with its own cloud-based IDE, and adds Interactive Planning, Devin Search, and Devin Wiki.

    AIWhy it matters: The release adds planning, codebase search, and auto-generated wikis to Devin, showing how an agent can prepare work before executing it.

Dec 9, 2024

Dec 9, 2024Mon
  1. Cognition Blog (Devin, Windsurf)AI score67

    Cognition makes Devin generally available to engineering teams from $500 a month

    Cognition is making Devin generally available to engineering teams starting at $500 a month, with no seat limits and access to its Slack integration, IDE extension, and API. The post recommends starting with small frontend bugs, first-draft PRs for backlog tasks, and targeted refactors, and shares open-source PR sessions where Devin resolved issues for projects including Anthropic MCP, Zod, and nanoGPT.

    AIWhy it matters: The post shows concrete open-source PR examples and the tasks where Devin works best, helping teams judge where an autonomous coding agent fits their workflow.

Sep 11, 2024

Sep 11, 2024Wed
  1. Cognition Blog (Devin, Windsurf)AI score60

    Cognition tests OpenAI o1 models in Devin's coding agent benchmark

    Cognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.

    AIWhy it matters: The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.

Mar 14, 2024

Mar 14, 2024Thu
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition reports Devin resolves 13.86% of SWE-bench issues end to end

    Cognition reports that its agent Devin resolved 79 of 570 sampled SWE-bench issues, a 13.86% success rate, without being given the files to edit. The report says this exceeds the best previous unassisted baseline of 1.96% and the best assisted result of 4.80%. It also describes the adapted evaluation setup, a 45-minute runtime limit, and cases where Devin failed on multi-file edits.

    AIWhy it matters: The report explains how SWE-bench was adapted for end-to-end agent evaluation, with failure cases that clarify where the 13.86% result comes from and its limits.

Mar 11, 2024

Mar 11, 2024Mon
  1. Cognition Blog (Devin, Windsurf)AI score88

    Cognition introduces Devin, an AI agent that works on software engineering tasks

    Cognition introduces Devin as an AI software engineer that can plan and execute complex engineering tasks with a shell, code editor, and browser. On SWE-bench, Devin resolved 13.86% of issues end-to-end, versus a previous state-of-the-art of 1.96%, on a random 25% subset of the dataset. Devin is in early access, with access available through a waitlist.

    AIWhy it matters: The post pairs Devin's end-to-end task demos with SWE-bench results against prior models, letting readers weigh the claimed capability against the evaluation setup.