Skip to contentSkip to stories

Updated

#Coding

Jun 1

Jun 1Mon
  1. Cognition Blog (Devin, Windsurf)AI score50

    Cognition launches Devin Desktop, the next generation of Windsurf

    AICognition has announced Devin Desktop, the next generation of Windsurf, which makes the Agent Command Center the default IDE surface for managing local and cloud agents, PRs, and context. Spaces let related agents share context, and Agent Client Protocol (ACP) support lets any ACP-compatible agent run alongside Devin. The IDE remains fully backwards-compatible with Windsurf, including editor extensions, keybindings, LSPs, and terminal workflows.

May 31

May 31Sun
  1. MiniMax BlogAI score82

    MiniMax M3 releases with 1M context, native multimodality and sparse attention

    AIMiniMax released M3, an open-weight model with a 1M-token context window, native image and video input, and desktop operation support. The post credits a new sparse attention architecture, MSA, for long-context gains, reporting over 9x prefilling and over 15x decoding speedups and 59.0% on SWE-Bench Pro. The API and MiniMax Code are available now, with the technical report and open weights promised within 10 days.

    Why it matters: The post pairs a new sparse attention design with benchmark figures and a 1M-token context window, letting readers judge the architecture's practical effect on long-context work.

May 28

May 28Thu
  1. Cognition Blog (Devin, Windsurf)AI score62

    Devin Tests Its Own Code Changes in the Cloud and Returns Proof

    AICognition describes autonomous testing in Devin, where the agent writes a source-grounded test plan, operates the app through computer use, and returns labeled screenshots and an annotated video. Login steps are handled by a deterministic testing skill, and the company says test runs approved per day more than doubled in recent months. Known limits include timing errors with transient UI elements and models sometimes triggering states through JavaScript instead of clicking the interface.

    Why it matters: The post explains how computer use, test plans, deterministic login scripts, and annotated recordings let Devin verify its own code changes end to end.

May 26

May 26Tue

May 22

May 22Fri
  1. AI Snake OilAI score60

    Google's $916 agent-built operating system claim lacks key methodology details

    AIGoogle claimed a team of agents built an operating system from a single prompt for about $916 in API fees, using Gemini 3.5 Flash and Antigravity 2.0. The authors argue the prompt was many thousands of lines, the scaffold and human intervention are undefined, and no code, logs, or similarity analysis were released to verify the claim. They still see value in such open-world evaluations, which need stronger methodological norms and independent scrutiny.

May 21

May 21Thu
  1. Tri DaoAI score44

    Transformers reduce to GEMM-plus-epilogue, enabling LLM-written fast kernels

    AITri Dao says that after a mathematical rewrite, all transformer operations can be expressed as a series of GEMMs with epilogues. Given a few optimized primitives, LLMs and novice humans can write near speed-of-light kernels for transformer ops. The related CODA work fuses memory-bound surrounding ops into the matmul epilogue, and LLMs can also write CODA kernels approaching speed-of-light.

May 20

May 20Wed
  1. Cognition Blog (Devin, Windsurf)AI score46

    Devin Gains Native Windows Environment for Building, Testing, and Migrating Apps

    AICognition's Devin AI software engineer can now build, run, and test code natively in its own Windows virtual machine, including migrating .NET Framework apps to .NET Core. The Windows capability is in beta for Enterprise Cloud and Dedicated Deployment customers, with the same SOC 2 Type II and ISO 27001 controls as the Linux version. Citi and Mercedes-Benz are named as existing Devin users.

May 19

May 19Tue
  1. koray kavukcuogluAI score72

    Google's Gemini 3.5 Flash beats Gemini 3.1 Pro on coding and agentic benchmarks

    AIGoogle's Gemini 3.5 Flash outperforms Gemini 3.1 Pro on Terminal-Bench 2.1 (76.2%), GDPval-AA (1656 Elo), and MCP Atlas (83.6%). The post also claims it is 4x faster than other frontier models, or 12x in Antigravity, and reports 83.6% on MMMU-Pro for multimodal performance.

    Why it matters: The post gives specific benchmark scores against Gemini 3.1 Pro, letting readers compare coding, agentic, and multimodal results directly.

May 18

May 18Mon
  1. Michael TruellAI score40

    Cursor's Composer 2.5 is a significant upgrade over Composer 2

    AIMichael Truell of Cursor says Composer 2.5 is a significant step up from Composer 2. He adds that this is only the start of work with SpaceXAI, with more improvements expected soon. Cursor's announcement describes the model as more intelligent, better at long-running tasks, and more reliable at complex instructions, with doubled included usage for the next week.

May 12

May 12Tue
  1. Cognition Blog (Devin, Windsurf)AI score40

    Devin now supports Android emulators for building and testing apps

    AICognition's Devin can now spin up an Android Virtual Device, letting it build, run, and test Android applications directly on its own machine. The new emulator support gives Devin an Android equivalent of computer and browser use, so it can open apps, inspect behavior, reproduce issues, and verify changes. The feature is available now for teams using Devin.

May 11

May 11Mon

Apr 30

Apr 30Thu
  1. OpenAI Alignment Research BlogAI score79

    OpenAI's Auto-review lets Codex agents act without constant human approval

    AIOpenAI released Auto-review in Codex, which replaces user approval at the sandbox boundary with a separate agent that approves or denies boundary-crossing actions. In internal deployment, Codex sessions stopped for human approval about 200x less often than in manual mode, and Auto-review approved around 99% of escalated actions. The post also states that Auto-review is not a guarantee of security and cannot protect against model scheming.

    Why it matters: The post explains how Auto-review replaces human approval at the sandbox boundary, with internal deployment figures and stated limits that help readers judge the tradeoff for coding agents.

  2. Andrej KarpathyAI score66

    Karpathy on agentic engineering, Software 3.0, and jagged AI capability

    AIAndrej Karpathy describes a December 2025 shift in which coding agents began producing larger, more reliable chunks of work, changing programming toward orchestrating agents. He argues that models automate what can be verified and that their capability is jagged, depending on verifiability and what labs emphasize in training, so users need to stay in the loop. He also says hiring, founder opportunities, and agent-native infrastructure should adapt to this shift.

Apr 27

Apr 27Mon
  1. Cognition Blog (Devin, Windsurf)AI score38

    Mercedes-Benz Deploys Devin and Windsurf Across Global Engineering Teams

    AIMercedes-Benz is deploying Cognition's Devin and Windsurf across its global engineering teams, from the United States to Europe and Asia. In a four-week pilot, Devin analyzed over 200,000 lines of COBOL code and cut modernization time from an estimated eight months to eight days. The company is now rolling out the full suite, with Windsurf for development, Devin as an autonomous cloud agent, and Devin for Terminal for the most complex tasks.

Apr 26

Apr 26Sun
  1. Xiaomi MiMoAI score87

    Xiaomi releases open-source MiMo-V2.5-Pro for long-horizon agentic coding

    AIXiaomi released and open-sourced MiMo-V2.5-Pro, a 1.02T-parameter Mixture-of-Experts model with 42B active parameters and a 1M-token context window. The company reports gains in agentic tasks, software engineering, and long-horizon work, including a Rust SysY compiler task finished in 4.3 hours across 672 tool calls. Weights and tokenizer are on Hugging Face, and API pricing is unchanged.

    Why it matters: The release pairs a 1.02T-parameter open-weight model with long-horizon agent results and token-efficiency claims, useful for judging its fit in coding and agent workflows.

Apr 25

Apr 25Sat

Apr 24

Apr 24Fri

Apr 23

Apr 23Thu
  1. Apple · new models on Hugging FaceAI score40

    Apple releases CADD-Base-7B, a masked diffusion model for code generation

    AIApple has released CADD-Base-7B on Hugging Face, a 7B masked diffusion language model for code generation that uses Continuously Augmented Discrete Diffusion (CADD) to guide discrete denoising with a continuous flow-matching signal. The model loads through Transformers with trust_remote_code, and its diffusion_generate method supports CADD sampling modes "weighted" and "argmax" with alg options such as "entropy" and "maskgit_plus". The release builds on DiffuCoder and reuses Dream's modeling architecture and generation utilities.

Apr 21

Apr 21Tue
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition says multi-agent systems work when only one agent writes

    AICognition reports that multi-agent setups work best when writes stay single-threaded and extra agents contribute intelligence instead of actions. It describes a code-review loop where a clean-context review agent catches bugs in Devin-written PRs, averaging 2 bugs per PR with roughly 58% severe. The post also says the smart-friend pattern, pairing a smaller primary model with a stronger one, has not yet worked well with asymmetrically weaker primaries and is an open training problem.

    Why it matters: The post gives concrete findings on which multi-agent setups work, including clean-context code review and smart-friend escalation, and where they still fail.

  2. Michael TruellAI score62

    Cursor partners with SpaceX to scale up Composer, with an option to acquire

    AICursor's Michael Truell says the company is partnering with the SpaceX team to scale up Composer, calling it a meaningful step toward building the best place to code with AI. The quoted SpaceX post says Cursor gives SpaceX the right to acquire Cursor later this year for $60 billion, or pay $10 billion for the work together. It also cites SpaceX's Colossus training supercomputer, described as a million H100-equivalent system, as a source of training capacity.

Apr 14

Apr 14Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score78

    Moonshot AI releases open-source Kimi K2.6 multimodal agentic model

    AIMoonshot AI released Kimi K2.6, an open-source native multimodal agentic model with 1T total and 32B activated parameters and a 256K context length. The model card reports benchmark results against GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro across agentic, coding, reasoning, and vision tasks, and supports swarms of up to 300 sub-agents.

    Why it matters: The model card gives specific agent swarm scale, context length, and benchmark comparisons against several frontier models, useful for judging its coding and agent capabilities.

Apr 13

Apr 13Mon
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition introduces SWE-check, a fast RL-trained bug detection model for Windsurf

    AICognition and Applied Compute RL-trained SWE-check, a specialized bug detection model for the Windsurf IDE. It matches frontier performance on in-distribution evals and is an order of magnitude faster with cheaper inference, though it trails frontier models on out-of-distribution evals (delta F1 0.29 versus 0.49 before training). A preview is available in Windsurf Next, with a mainstream release planned.

    Why it matters: The post explains how production environment replication, reward linearization, and two-phase post-training trade bug-detection quality against latency for an IDE specialist model.

  2. Cognition Blog (Devin, Windsurf)AI score49

    Devin Introduces New Self-Serve Plans and Charges for Ask Devin and Devin Review

    AIDevin is retiring its Core and Team plans for a new lineup of Free, Pro at $20/month, Max at $200/month, Teams with usage-based billing and an $80/month minimum, and custom-priced Enterprise. Ask Devin's Deep Mode, Devin Review after a 2-week free trial, and higher-quality DeepWiki generation will move to usage-based billing, with DeepWiki's existing generation and open-source Devin Review remaining free. Self-serve usage beyond included quota will be billed in dollars rather than ACUs.

Apr 9

Apr 9Thu
  1. Andrej KarpathyAI score45

    Karpathy says AI capability gap stems from uneven use and training

    AIAndrej Karpathy argues that people judging AI from free-tier ChatGPT or Advanced Voice Mode miss the strong capabilities of current agentic models like OpenAI Codex and Claude Code. He says gains are "peaky," concentrated in verifiable technical domains like programming and math that suit reinforcement learning and attract B2B investment, while writing and everyday advice improve less. Those who use frontier agentic tools professionally in these fields see far greater capability, which is why the two groups talk past each other.

Apr 8

Apr 8Wed
  1. MiniMax · new models on Hugging FaceAI score78

    MiniMax releases open-weight MiniMax-M2.7 with agent and coding gains

    AIMiniMax has released MiniMax-M2.7 on Hugging Face, describing it as its first model to participate in its own evolution. The source reports 56.22% on SWE-Pro, 46.3% on Toolathon, and 62.7% on MM ClawBench, and says an internal version autonomously optimized a programming scaffold over 100+ rounds for a 30% performance improvement.

    Why it matters: The source ties its benchmark claims to a self-evolution process and a named comparison set, which helps readers weigh how the reported gains were achieved.

Apr 7

Apr 7Tue
  1. Cognition Blog (Devin, Windsurf)AI score70

    How Devin Is Modernizing COBOL at Fortune 500 Companies

    AICognition describes how Devin handles COBOL modernization at several Fortune 500 companies, citing a shortage of COBOL developers and 68% failure rates for such efforts. The post identifies three obstacles for agents: untraceable data across copybooks, little COBOL in model training, and no way to run code on Linux-based VMs. It says Devin succeeds on documentation, batch migrations, and large-scale refactoring, while transactional workloads remain out of reach.

    Why it matters: The post explains why agents struggle with COBOL and which workloads they can migrate, giving a framework for judging where automation fits legacy systems.

Apr 6

Apr 6Mon
  1. Z.ai Release NotesAI score34

    Z.ai's GLM-5.3 and GLM-5.2 Lead Open-Source Coding and Long-Context Models

    AIZ.ai's GLM-5.3 delivers a 50% coding gain over GLM-5.2 on Z.ai Code Bench, reaching open-source state-of-the-art on public benchmarks including Terminal Bench 3.0. GLM-5.3-Flash uses 320B total parameters with 18B activated, combining linear and sparse attention to reduce compute and KV-cache needs. GLM-5.2 supports a 1M lossless context window for long-horizon tasks.

  2. Cognition Blog (Devin, Windsurf)AI score44

    Windsurf releases SWE-1.6, a software engineering model optimized for speed and user experience

    AIWindsurf has made SWE-1.6, its model for software engineering agents, generally available, with the company saying it improves on the SWE-1.6 Preview by reducing overthinking, looping, and sequential tool calls. The model is free for three months, with a free version offered at 200 tok/s through Fireworks and a faster paid version at 950 tok/s through Cerebras.

Apr 3

Apr 3Fri
  1. Z.ai (GLM) · new models on Hugging FaceAI score73

    Z.ai releases GLM-5.1, a flagship model for agentic engineering

    AIZ.ai has released GLM-5.1, its next-generation flagship model for agentic engineering, with stronger coding than GLM-5. The model is described as staying effective over longer agentic tasks, sustaining optimization over hundreds of rounds and thousands of tool calls. The release lists benchmark results including SWE-Bench Pro at 58.4 and Terminal-Bench 2.0 at 63.5, and local deployment is supported through SGLang, vLLM, xLLM, Transformers, and KTransformers.

    Why it matters: The release gives benchmark tables against several rival models, letting readers compare GLM-5.1's coding and agentic results with GLM-5 and frontier systems.

Apr 2

Apr 2Thu
  1. AI Futures ProjectAI score62

    AI Futures Project shortens Automated Coder timelines to mid 2028

    AIAI Futures Project moved Daniel Kokotajlo's Automated Coder median from late 2029 to mid 2028 and Eli's from early 2032 to mid 2030. The main reasons cited are a faster METR time horizon doubling time and the impressive results of Claude Opus 4.6. The authors also say progress in agentic coding has been faster than expected over the past 3 to 5 months.

Mar 31

Mar 31Tue
  1. Mistral AI · new models on Hugging FaceAI score76

    Mistral Medium 3.5 releases as a 128B dense merged model with vision

    AIMistral AI released Mistral Medium 3.5, a dense 128B model with a 256k context window that handles instruction-following, reasoning, and coding in a single set of weights. It replaces Mistral Medium 3.1, Magistral, and Devstral 2, and reasoning effort is configurable per request. The model accepts text and image input and is released under a Modified MIT License that excludes companies with large revenue.

    Why it matters: The release merges instruction, reasoning, and coding into one 128B model with per-request reasoning control, giving developers one set of weights to compare against separate specialized models.