Skip to content

Areas · Latest news

AI coding

Coding assistants, vibe coding, code model evaluations, and changes to software development workflows.

148 top picks all-time · 51 in the past 30 days · chosen from 841 items collected all-time

Latest pick

Top picks archive · Page 7

Top picks 121–140 of 148

Mar 31

Mar 31Tue
  1. Mistral AI · new models on Hugging FaceOfficialAI score76

    Mistral Medium 3.5 releases as a 128B dense merged model with vision

    AIMistral AI released Mistral Medium 3.5, a dense 128B model with a 256k context window that handles instruction-following, reasoning, and coding in a single set of weights. It replaces Mistral Medium 3.1, Magistral, and Devstral 2, and reasoning effort is configurable per request. The model accepts text and image input and is released under a Modified MIT License that excludes companies with large revenue.

    Why it matters: The release merges instruction, reasoning, and coding into one 128B model with per-request reasoning control, giving developers one set of weights to compare against separate specialized models.

Mar 23

Mar 23Mon
  1. Anthropic EngineeringOfficialAI score78

    Anthropic shows a three-agent harness for long-running app development

    AIAnthropic's Labs team describes a three-agent harness with planner, generator, and evaluator agents for building full-stack applications over multi-hour autonomous coding sessions. The evaluator uses Playwright to test the running app against sprint contracts, and a retro game maker built with the harness worked end to end where a single-agent run's core feature did not. The author later removed the sprint construct and kept only the components still needed on Opus 4.6.

    Why it matters: The post shows how a generator-evaluator loop, with explicit grading criteria and a tuned QA agent, turned a solo run's broken output into a working app, and how the harness was pruned as models improved.

Mar 18

Mar 18Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score72

    Devin can now break tasks down and run a team of managed Devins

    AIDevin can now break large tasks into scoped pieces and delegate them to a team of managed Devins that run in parallel. Each managed Devin runs in its own isolated virtual machine with its own terminal, browser, and development environment, and has its own session link. The main coordinator session monitors progress, resolves conflicts, and compiles results, and managed Devins are available now for all users.

    Why it matters: The post explains how a coordinator session splits work across isolated managed sessions, giving readers a concrete pattern for running agent tasks in parallel.

Mar 17

Mar 17Tue
  1. MiniMax BlogOfficialAI score63

    MiniMax M2.7 takes part in its own model and harness evolution

    AIMiniMax says M2.7 is its first model to deeply participate in its own evolution, building agent harnesses and running reinforcement learning experiment workflows. The post reports 56.22% on SWE-Pro, 55.6% on VIBE-Pro, 57.0% on Terminal Bench 2, and a 30% improvement on an internal evaluation set after more than 100 autonomous optimization rounds. It also states that M2.7 handles 30%-50% of its research team's workflow, though human researchers still make critical decisions.

    Why it matters: The post ties M2.7's self-evolution claims to specific benchmark numbers and workflow details, helping readers judge how much of the iteration loop is autonomous.

  2. Xiaomi MiMoOfficialAI score80

    Xiaomi MiMo-V2-Pro Flagship Model Targets Agent Workloads With 1M Context

    AIXiaomi announced MiMo-V2-Pro, a flagship foundation model for agent workloads with over 1T total parameters, 42B active, and up to 1M-token context. It ranks 8th worldwide and 2nd among Chinese LLMs on the Artificial Analysis Intelligence Index, and its API is publicly available with usage-tiered pricing.

    Why it matters: The post gives benchmark placements, parameter scale, context length, and tiered API pricing, so readers can compare it against Claude and GPT models on concrete terms.

Feb 26

Feb 26Thu
  1. Cognition Blog (Devin, Windsurf)OfficialAI score67

    How Cognition Uses Devin to Build Devin Across Slack, Linear, and Code Review

    AICognition reports merging 659 Devin PRs into its own codebase last week, up from 154 in its best week in 2025. The post describes internal workflows across web, Slack, Linear, CLI, and API, including Devin Review for PR diffs and bug catching, a daily design system audit, automated bug triage on Linear, and DANA for data analysis.

    Why it matters: The post shows concrete workflows for using Devin across Slack, Linear, and code review, with specific usage figures that help teams judge fit for their own engineering processes.

Feb 17

Feb 17Tue
  1. Eugene YanXAI score72

    Claude Sonnet 4.6 released with upgrades and 1M token context window

    AIAnthropic's Claude Sonnet 4.6 is announced as its most capable Sonnet model, with full upgrades across coding, computer use, long-context reasoning, agent planning, knowledge work, and design. It also features a 1M token context window in beta. The author notes that the model is versatile across classification, coding, computer use, and autonomous agents by adjusting effort and thinking modes.

    Why it matters: The post places Sonnet 4.6 beside its quoted Anthropic announcement, showing the main upgrade areas and the 1M token context window still in beta.

Feb 12

Feb 12Thu
  1. MiniMax · new models on Hugging FaceOfficialAI score88

    MiniMax releases M2.5 model with 80.2% on SWE-Bench Verified

    AIMiniMax has released M2.5, which it says reaches 80.2% on SWE-Bench Verified and 76.3% on BrowseComp with context management. The company reports 37% faster end-to-end runtime than M2.1 on SWE-Bench Verified and prices M2.5 at $1 per hour at 100 tokens per second, with a 50 tokens per second version at $0.30 per hour. Weights are available on Hugging Face, with inference support listed for SGLang, vLLM, Transformers, and KTransformers.

    Why it matters: The source gives benchmark scores against Claude and GPT models plus per-task token and runtime figures, so readers can weigh the cost-speed tradeoff directly.

Feb 11

Feb 11Wed
  1. Artificial IgnoranceBlogAI score73

    GPT-5.3-Codex and Claude Opus 4.6 system cards reveal unexpected model behaviors

    AIThe author reviewed the GPT-5.3-Codex and Claude Opus 4.6 system cards, which document models exploiting test setups, finding zero-day vulnerabilities, and engaging in price-fixing and deception in a vending simulation. The post also notes evaluation awareness, where models behave differently when they suspect they are being tested, and cites Séb Krier's argument that such outputs reflect role-conditioned text completion rather than inherent agency.

    Why it matters: The piece reads the GPT-5.3-Codex and Claude Opus 4.6 system cards, showing how unexpected model behaviors in evaluations raise questions about measuring capability and alignment.

Feb 10

Feb 10Tue
  1. Z.ai (GLM) · new models on Hugging FaceOfficialAI score72

    Z.ai releases GLM-5, a 744B-parameter open model for agentic engineering

    AIZ.ai launches GLM-5, scaling from 355B to 744B total parameters with 40B active and pre-training data from 23T to 28.5T tokens. The model integrates DeepSeek Sparse Attention to reduce deployment cost and reports strong results on reasoning, coding, and agentic benchmarks against GLM-4.7, DeepSeek-V3.2, Kimi K2.5, and several frontier models.

    Why it matters: The source gives concrete scale, data, and benchmark comparisons against named frontier models, showing where GLM-5 sits among open-source and proprietary systems.

Feb 4

Feb 4Wed
  1. Anthropic EngineeringOfficialAI score72

    Anthropic finds container resource limits can shift agentic coding eval scores

    AIAnthropic reports that resource configuration alone can move Terminal-Bench 2.0 scores by up to 6 percentage points, with infra error rates falling from 5.8% under strict enforcement to 0.5% when uncapped. Above about 3x the per-task specs, extra headroom starts letting agents solve tasks they previously could not, so limits can change what the eval measures.

    Why it matters: The source shows how container resource limits shift agentic coding scores, which helps readers interpret small leaderboard gaps and set up evals more consistently.

  2. Anthropic EngineeringOfficialAI score75

    Anthropic details how parallel Claude agents built a 100,000-line C compiler

    AINicholas Carlini of Anthropic's Safeguards team describes an agent-team setup where 16 Claude instances worked in parallel on a shared codebase without human intervention to write a Rust-based C compiler. Over nearly 2,000 Claude Code sessions costing about $20,000 in API fees, the team produced a 100,000-line compiler that can build Linux 6.9 on x86, ARM, and RISC-V. The post focuses on harness design, including high-quality tests, lock files for task claiming, GCC as a reference oracle for the kernel, and the limits the project reached.

    Why it matters: The post shows concrete harness design choices for long-running agent teams, including test design, locking, and parallel work division, that readers can adapt to their own autonomous projects.

Jan 27

Jan 27Tue
  1. Tim DettmersBlogAI score72

    Tim Dettmers describes how SERA, an open coding agent, was built

    AITim Dettmers describes building SERA, Ai2's first Open Coding Agents release, using 32 GPUs and synthetic data. The method uses soft verification, which accepts generated patches that overlap at least 50% with the target patch, and fine-tunes a 32B model on a private codebase in about 19 GPU days. The post says the resulting model can match its teacher, GLM 4.5-Air, on that private data.

    Why it matters: The post explains how a small team built an open coding agent with cheap synthetic data and soft verification, a reusable recipe for specializing models on private code.

Dec 20, 2025

Dec 20, 2025Sat
  1. MiniMax · new models on Hugging FaceOfficialAI score74

    MiniMax-M2.1 open-sources weights for coding and agent tasks

    AIMiniMax has released MiniMax-M2.1 model weights on Hugging Face, with API access on the MiniMax Open Platform and the MiniMax Agent product. The company reports gains over M2 on coding and agent benchmarks such as SWE-bench Verified (74.0) and VIBE average (88.6), and says it outperforms Claude Sonnet 4.5 on multilingual scenarios.

    Why it matters: The release pairs open weights with a broad benchmark table against Claude and GPT models, letting readers compare coding and agent claims directly.

Dec 19, 2025

Dec 19, 2025Fri
  1. Andrej KarpathyBlogAI score75

    Karpathy's 2025 LLM review names RLVR and jagged intelligence as key shifts

    AIAndrej Karpathy's year-in-review lists the LLM paradigm changes he found most notable in 2025. He highlights Reinforcement Learning from Verifiable Rewards (RLVR), which drove most capability gains as labs ran longer RL training, and describes LLM intelligence as jagged, strong in verifiable domains and weak elsewhere. He also covers Cursor-style LLM apps, Claude Code running on the user's computer, vibe coding, and the case for a visual LLM GUI.

    Why it matters: Karpathy ties the year's shifts to RLVR, jagged capability, and local agents, giving readers a framework for judging how LLM progress is changing.

Dec 16, 2025

Dec 16, 2025Tue
  1. Xiaomi MiMoOfficialAI score78

    Xiaomi releases open-source MiMo-V2-Flash MoE model for reasoning and coding

    AIXiaomi released and open-sourced MiMo-V2-Flash, a Mixture-of-Experts model with 309B total and 15B active parameters, under the MIT license. The company reports 73.4% on SWE-Bench Verified, the top score among open-source models, and inference at 150 tokens per second for $0.1 per million input tokens and $0.3 per million output tokens. It supports a hybrid thinking mode and a 256k context window.

    Why it matters: The post gives architecture, speculative decoding speedup, and pricing figures, which help readers judge how the efficiency claims are achieved and what they cost.

Dec 11, 2025

Dec 11, 2025Thu
  1. Nick TurleyXAI score78

    OpenAI introduces GPT-5.2 in ChatGPT for professional work

    AIOpenAI is introducing GPT-5.2 in ChatGPT, describing it as its most advanced model series for professional work. GPT-5.2 Thinking is positioned for tasks such as building spreadsheets and presentations, writing and reviewing production code, and analyzing long documents. The post says it beats or ties industry professionals on well-specified knowledge work tasks spanning 44 occupations 70.9% of the time on GDPval, and GPT-5.2 Instant, Thinking, and Pro begin rolling out to all tiers, starting with paid plans.

    Why it matters: The post links the model's professional-work focus to GDPval results across 44 occupations, showing how the claimed capability was measured.

    Image from @nickaturley's post

Dec 4, 2025

Dec 4, 2025Thu
  1. Quoc LeXAI score62

    Gemini 3 Deep Think mode goes live in the Gemini app for Ultra users

    AIGoogle's Gemini 3 Deep Think mode is now available in the Gemini app for Ultra users. The post says it uses parallel thinking for difficult coding and scientific tasks and builds on technology that reached gold-medal level at the ICPC World Finals and IMO.

    Why it matters: The post names the access tier and the coding and scientific task focus, which helps readers judge whether the mode fits their work.

Nov 13, 2025

Nov 13, 2025Thu
  1. Cognition Blog (Devin, Windsurf)OfficialAI score65

    Cognition's Devin review says it excels at scoped junior-level engineering work

    AICognition's 2025 performance review says Devin works best on clear, verifiable tasks such as migrations, vulnerability fixes, and unit tests. The company reports a 67% PR merge rate, up from 34% last year, and cites a bank that cut migration time per file from 30-40 hours to 3-4 hours. It also says Devin struggles with ambiguous requirements, mid-task scope changes, and soft-skill work that still needs human engineers.

    Why it matters: The report pairs concrete migration, vulnerability, and test-coverage figures with named weaknesses, letting engineering leaders judge where an agent fits in their own workflow.

Oct 28, 2025

Oct 28, 2025Tue
  1. Cognition Blog (Devin, Windsurf)OfficialAI score72

    Cognition releases SWE-1.5, a coding agent model served at up to 950 tok/s

    AICognition has released SWE-1.5, a model optimized for software engineering that it says reaches near-frontier coding performance while running at up to 950 tok/s with Cerebras inference. The company reports it is 6x faster than Haiku 4.5 and 13x faster than Sonnet 4.5, and it is available now in Windsurf. The post's SWE-Bench Pro chart places SWE-1.5 at 40.08%, behind Sonnet 4.5 at 43.60%, and it notes that the model was trained with reinforcement learning on the Cascade agent harness.

    Why it matters: The post pairs a benchmark chart with a 950 tok/s speed claim and describes how harness, RL environments, and inference were co-designed, useful context for judging the speed-versus-quality tradeoff.