Skip to contentSkip to stories

Updated

#Coding

Sep 13

Sep 13Sun
  1. Fireworks AI BlogAI score52

    Fireworks adds DeepSeek-V4.1-Flash, matching GPT-6 Astra coding accuracy at 1/15th the cost

    AIFireworks AI reports that DeepSeek-V4.1-Flash scores 74.34% pass@1 on DeepSWE at $0.430 per task, close to GPT-6-Astra's 74.12% at $6.524. On Terminal-Bench 2.1 it scores 86.5% against Astra's 87.5% at about 12x lower cost per task, while on HLE it trails Astra alone at 34.52% versus 50.40%. The post also reports that a combined oracle router reaches 54.80% on HLE, and that serverless and dedicated API access is available with US-hosted endpoints coming soon.

Sep 11

Sep 11Fri
  1. Augment Code BlogAI score80

    Augment Code details how its software factory raised output per developer 4.5×

    AIAugment Code reports that size-adjusted output per active developer rose from 12.3 to 55.7 between November 2025 and July 2026, while median time to merge fell from 11.2 to 3.1 hours. The post says the company added specialized agents wherever work was piling up, across planning, review, verification, feedback, and incident response, and kept engineers responsible for product decisions, architecture, and production risk.

    Why it matters: The post pairs internal productivity and quality metrics with the order in which agents were added, showing how review and verification bottlenecks shaped a software delivery pipeline.

  2. Baseten BlogAI score62

    DeepSeek-V4.1-Flash arrives on Baseten with a split prefill architecture

    AIDeepSeek released open weights for V4.1-Flash, which Baseten now offers through its Model APIs. The model has 552B total parameters, 8B active for prefill and 16B for decode, a 1M token context window, and text plus image input. Its Causal Encoder-Decoder design runs only the encoder during prefill and reuses a projected KV cache, and the source reports the global KV cache at a quarter of V4-Flash's memory.

    Why it matters: The post explains how the CED architecture splits prefill and decode compute and cuts KV cache memory, which matters for coding agent costs.

  3. Cognition Blog (Devin, Windsurf)AI score51

    Cognition introduces Fusion in Devin Desktop and CLI for lower-cost coding

    AICognition is making Fusion available in Devin Desktop and CLI, a harness where a frontier lead model plans and reviews while a cheaper sidekick executes. Across listed coding benchmarks, Cognition reports Fusion cuts cost per task by about 11% to 46% versus the lead model alone, while the sidekick does the implementation work. The post recommends pairing Fable 5.1 with SWE-2, and argues price per task matters more than price per token.

  4. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score72

    Shanghai AI Lab releases Atria Dawn Preview, an agentic model built on GLM-5.2

    AIShanghai Artificial Intelligence Laboratory has released Atria Dawn Preview, an agentic model built on the 744B-parameter MoE GLM-5.2 foundation model, with a 256K context window. The release page reports benchmark results across search, coding, tool use, productivity, and cybersecurity, and describes text-only setup for Codex and Claude Code.

    Why it matters: The release page gives a full benchmark table against named rivals and setup steps for Codex and Claude Code, useful for anyone evaluating agentic models.

Sep 10

Sep 10Thu
  1. Cognition Blog (Devin, Windsurf)AI score66

    Cognition releases SWE-2, a coding model trained with cost-penalized RL

    AICognition introduces SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while costing 64% less. The post attributes the gains to an RL algorithm that trains all reasoning-effort levels in one run, with cost penalties tuned to the base model's Pareto frontier. SWE-2 is available starting today in Devin Desktop and CLI, with rollout to Devin Web and Fusion.

    Why it matters: The post explains how the cost penalty and length-weighted baseline are derived, which helps readers judge the tradeoffs in coding model post-training.

Sep 9

Sep 9Wed
  1. Mistral AIAI score54

    Mistral details how AI agents migrated 40,000 lines of Fortran to C++

    AIMistral AI helped a European energy operator migrate 40,000 lines of Fortran 77 to C++ for a reservoir simulator with no test suite. The post explains a parity harness that checks numerical agreement between the two codebases, and a workflow where agents coder, tester, and reviewer migrate modules under human review. Its authors note the approach covered the self-contained first sprint of 40,000 of 300,000 lines and that dependent systems would bring additional challenges.

Sep 8

Sep 8Tue
  1. Google Developers BlogAI score36

    Google Developers Blog outlines behavioral evals for guarding AI coding agents against regressions

    AIGoogle Developers Blog argues that teams building AI coding agents should replace end-to-end benchmark scores with behavioral evaluations that test discrete, observable actions. Examples include asking clarifying questions on underspecified prompts, running a local validator before marking a build change complete, and consulting live search for current information. The post recommends fast, deterministic unit-style checks, outcome-based LLM-as-a-judge checks for complex tasks, and batch runs that track aggregate pass rates over time.

Sep 6

Sep 6Sun
  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score62

    OpenBMB releases MiniCPM5-2B, a 2B open-source model with open training data

    AIOpenBMB has released MiniCPM5-2B, a dense 2B Transformer built for on-device and resource-constrained deployment, with an average score of 53.9 in its comparison set. The release also opens the UltraData datasets behind it, including UltraX, UltraData-Code, UltraData-SFT-Agent-2609 and UltraData-RL-2609, and includes GGUF, MLX, GPTQ and DSpark variants for common runtimes.

    Why it matters: The release pairs a 2B model with open training datasets and reports per-benchmark comparisons against named same-size and larger models, letting readers check the claims directly.

Sep 2

Sep 2Wed

Sep 1

Sep 1Tue
  1. Anthropic · YouTubeAI score78

    Anthropic releases Claude Fable 5.1, an upgrade to its most capable model class

    AIAnthropic has released Claude Fable 5.1, the latest upgrade to its most capable class of models, and it is available everywhere today. The company says it handles complex, long-running, multi-step work and avoids shortcuts when fixing root causes of software issues. At lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost, according to Anthropic's benchmarks.

    Why it matters: The source names the upgraded model class and its cost tradeoff at lower effort levels, which helps readers weigh it against the earlier version for their own workloads.

Aug 31

Aug 31Mon
  1. Claude Apps Release NotesAI score72

    Anthropic launches Claude Fable 5.1 and Claude Mythos 5.1 models

    AIAnthropic has launched Claude Fable 5.1 and Claude Mythos 5.1, which it describes as the world's most advanced models for coding and knowledge work. The release notes link to a blog post with more details, but the notes themselves give no benchmarks or specifications.

    Why it matters: The source names two new model versions and points to a companion blog post, so readers can compare the release details there.

Aug 27

Aug 27Thu
  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score65

    OpenBMB releases MiniCPM5-2B-SFT, a 2B open model with SFT-only checkpoint

    AIOpenBMB released MiniCPM5-2B-SFT, an SFT-only BF16 checkpoint taken before RL and OPD, within its MiniCPM5-2B series. The model is a 2B dense Transformer built for on-device and local deployment, with 131,072-token context and the same training recipe as the final release.

    Why it matters: The source gives concrete benchmark averages against same-size and larger models, plus released training data and multiple deployment formats, useful for judging a compact on-device model.

Aug 26

Aug 26Wed
  1. Cursor ChangelogAI score46

    Cursor Cloud Agents now let you start projects from scratch without a repo

    AICursor Cloud Agents no longer require a connected GitHub or other third-party SCM provider to begin work. Users select "Start from scratch" in the repo picker, and Cursor creates an Origin repo in the background that can be saved as a private or internal repo via "Create repo." Cursor also now port-forwards the cloud agent's live environment to the browser for previews, and a connected Vercel account lets users publish a live URL.

Aug 25

Aug 25Tue
  1. Fireworks AI BlogAI score40

    DeepSeek V4 Pro 0813 Tops SWE-Bench and Cuts Cost per Solved Task

    AIDeepSeek V4 Pro 0813 scored 95.2% on SWE-Bench Verified, ahead of Kimi K3 at 92.6% and Fable 5 at 85.4%, in Fireworks AI's eval runs. It costs $0.309 per solved task on SWE-bench versus $0.808 for Fable 5, and it is available through Fireworks serverless and dedicated endpoints, with SFT, DPO, and RFT training support. Its 1M-token context window and native tool calling target long-horizon agentic workloads, though its Java accuracy on Aider Polyglot (48.9%) trails Fable 5 (74.5%).

  2. Fireworks AI BlogAI score46

    DeepSeek V4 Pro Solves Security Tasks at Half the Cost Per Success

    AIDeepSeek V4 Pro 0813 recorded zero refusals across 840 adversarial security tasks in CyberGym testing, solving them at about half the cost per success of the top-scoring model tested, Kimi K3. In the 697-task common cohort, V4 Pro reached a 53.7% reward rate at $2.50 per solved task, versus 47.6% and $9.64 for GPT-5.5 and 5.9% and $33.28 for Claude Opus 4.8.

  3. Z.ai Release NotesAI score62

    Z.ai releases GLM-5.3-Flash with native visual capabilities and hybrid architecture

    AIZ.ai has released GLM-5.3-Flash, a model with native visual capabilities that observe interfaces, rendering results, and interaction feedback across code, browsers, and GUIs. It uses a hybrid linear and sparse attention architecture with 320B total parameters and 18B activated, which the company says significantly reduces compute and KV-cache requirements. The release notes also describe support for office document and financial research workflows.

    Why it matters: The release notes give GLM-5.3-Flash's architecture, parameter counts, and cybersecurity findings, which make the model's scope concrete for comparison with earlier GLM releases.

  4. Z.ai (GLM) · new models on Hugging FaceAI score72

    Z.ai releases GLM-5.3 open weights with gains from post-training

    AIZ.ai released GLM-5.3 on Hugging Face, built on the same base model as GLM-5.2, with all gains coming from post-training. The source reports a 50% improvement over GLM-5.2 on Z.ai Code Bench and open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam, with a benchmark table comparing it against Kimi K3, DeepSeek-V4 Pro-0813, Qwen3.8-Max, and others.

    Why it matters: The source gives benchmark tables against GLM-5.2 and rival models, showing where the post-training gains concentrate in coding and cyber tasks.

Aug 18

Aug 18Tue
  1. Cursor ChangelogAI score62

    Cursor adds event subscriptions, custom modes, and subagent VMs for cloud agents

    AICursor's update lets cloud agents subscribe to PRs, Slack threads, and scheduled tasks, and wake when something happens. It also adds custom modes that pin a skill in chat, subagents that run on their own virtual machines, and a /goal command for long-lived objectives. Users can also send steering messages while an agent works, with follow-ups applied at the next tool call.

    Why it matters: The release lists concrete agent controls such as event subscriptions, custom modes, subagent VMs, and /goal, showing how cloud agents may run longer tasks with less manual steering.

  2. Replit BlogAI score46

    Replit launches Free Mode, letting subscribers build 30x more with Agent for $20 per month

    AIReplit has launched Free Mode, a new way to use its Agent that lets Core subscribers create up to 30x more on their monthly subscription, with everyday tasks no longer consuming credits. Free Mode is powered by OpenAI's GPT-5.6 Luna and is available to Core and Pro users until they reach usage limits that reset every 5 hours. Core subscribers also receive up to 30 hours per month of chat, and the company is offering the plan for $20 per month.

  3. Cursor BlogAI score68

    Cursor explains Continuity, a WAL-based Git storage system

    AICursor's blog describes Continuity, its Git storage system, which stores each push as a write-ahead log entry in S3-compatible object storage. The article contrasts this design with GitHub's earlier Spokes system, which used three-phase commit replication across local disks. Continuity uses stateless replicas that catch up from the log, and the article reports write throughput of up to 120 pushes/s on S3 Standard and over 300 pushes/s on S3 Express One Zone.

    Why it matters: The article explains why hosting Git at scale is hard and how Continuity's WAL-based design compares with the earlier Spokes approach, which is useful background for infrastructure work.

Aug 17

Aug 17Mon
  1. Z.ai Release NotesAI score63

    Z.ai releases GLM-5.3 with stronger coding and vulnerability discovery

    AIZ.ai's release notes announce GLM-5.3, which the company says delivers a 50% gain over GLM-5.2 on Z.ai Code Bench and reaches open-source SOTA on public benchmarks including Terminal Bench 3.0. The company also reports that GLM-5.3 matches Mythos 5 in white-box code review and vulnerability discovery, identifying 2,436 vulnerabilities in real-world targets, 1,097 of them medium- or high-severity. A separate GLM-5.3-Flash entry describes native visual capabilities and a hybrid architecture with 320B total and 18B activated parameters.

    Why it matters: The release notes show GLM-5.3's coding and cybersecurity gains, with a vulnerability count, letting readers compare it against Z.ai's prior GLM-5.x line and other coding models.

  2. Replit BlogAI score60

    Replit adds black-box pen tests that probe apps like external attackers

    AIReplit now offers black-box pen tests that scan deployed apps over the network and browser, with no access to source code. A Level 3 scan runs them alongside the existing white-box code scan, and the source notes the two catch different kinds of flaws.

    Why it matters: The post explains how black-box scans test an app like an outside attacker, showing why source-code review alone misses some exposed doors.

Aug 16

Aug 16Sun
  1. Cursor ChangelogAI score60

    Cursor launches Origin, a code hosting service with GitHub sync

    AICursor begins rolling out Origin, its code hosting feature, in early beta to all paid plans, excluding enterprise orgs whose admins opt out. Repos can be hosted on Origin, where Origin is the source of truth, or synced from GitHub, where GitHub stays the source of truth and pull requests sync both ways. Vercel, Depot, and Buildkite integrations are already available, and agent-native features are slated to ship soon.

    Why it matters: The source specifies how Origin hosts repos alongside GitHub sync, showing how the hosting source of truth differs between the two types of repo.

Aug 15

Aug 15Sat
  1. Prime Intellect BlogAI score73

    Prime Intellect tests frontier models on 153 autonomous nanoGPT research runs

    AIPrime Intellect ran 153 autonomous runs on the nanoGPT optimizer speedrun across 18 frontier models, with runs lasting up to eight days on 8xH200s. The results show a large gap between models at every stage of the research process, though none of the runs produced a fundamentally new method.

    Why it matters: The experiment measures how frontier models conduct autonomous research, showing large gaps between models in experiment choice, execution, and result interpretation.

Aug 14

Aug 14Fri
  1. Augment Code BlogAI score62

    Augment rebuilds its Auggie CLI harness on Pi, cutting SWE-bench Pro task cost 53%

    AIAugment rebuilt the Auggie CLI harness as v2, forking the open-source Pi coding harness and moving its context engine into Pi's extension system. On SWE-bench Pro at the same pass rate, Auggie v2 completes a task for $1.27 versus $2.70 for Claude Code, which is 53% cheaper. The gains come mainly from a narrower tool surface, one bash tool plus read, edit, and write, and from codebase retrieval that reduces exploration turns.

    Why it matters: The post traces the design trade-offs behind each harness choice and ties them to measured token and cost differences, useful for anyone weighing agent tool surfaces.

Aug 13

Aug 13Thu
  1. DeepSeek API NewsAI score62

    DeepSeek-V4-Pro Reaches GA with Agent Gains and Peak/Off-Peak API Pricing

    AIDeepSeek has made DeepSeek-V4-Pro generally available on its app, web, and API, with the API model name set to deepseek-v4-pro. The release reports agent benchmark results, including 87.9 on Terminal Bench 2.1 and 74.1 on Toolathlon-Verified. It also adds native OpenAI Responses API support, low/high/max thinking effort levels, and off-peak API prices set at half of peak prices starting 16:00 UTC on August 16, 2026.

    Why it matters: The update pairs new agent benchmark results with API format and pricing changes, so developers can judge both capability and cost impact before migrating.

Aug 12

Aug 12Wed
  1. Factory NewsAI score40

    Factory Launches Agent Effectiveness to Link Droid Usage to Delivery Outcomes

    AIFactory's Agent Effectiveness, now in Private Preview within Factory Analytics, connects Droid sessions to cycle time, work intent, and shipped artifacts drawn from project, issue-tracking, and source control tools. Its Throughput, Output, and Attribution views show where delivery is speeding up, how spend splits across feature, maintenance, bug-fixing, and exploration work, and which projects and issues the output maps to. Admins enable it by connecting Jira, Linear, GitHub, or GitLab and turning on the Advanced Analytics enterprise control.

  2. Cursor ChangelogAI score42

    Cursor Cloud Agents Start 3x Faster With Builds

    AICursor's Cloud Agents now start from prebuilt copies of development environments, cutting startup time by 3x, with environments booting 10x faster internally and 3x faster time to first token. Builds are included at no additional cost, and failed builds are not activated, so agents keep using the last successful build while users debug in the background.

Aug 11

Aug 11Tue
  1. Zed BlogAI score72

    Zed introduces Delta, a multiplayer environment for coding with agents

    AIand reviewing their code, and invites first users into a private beta. Delta keeps code and conversations connected through DeltaDB, which captures edits and conversations between git commits and works with existing repositories. The app also supports cloud runners, browser-based sharing, and live syncing of Claude Code sessions.

    Why it matters: The post explains how the new Delta app links conversations with code history, which clarifies a shift in how teams review agent-written changes.

Aug 7

Aug 7Fri
  1. Qwen · new models on Hugging FaceAI score88

    Qwen releases open-weight Qwen3.8-2.4T-A95B, a 2.4T-parameter MoE model

    AIQwen has released the Qwen3.8-2.4T-A95B model weights on Hugging Face, with 2.4T total and 95B activated parameters in a mixture-of-experts design. The release supports reasoning_effort levels and a 262,144-token native context extensible to 1,010,000 tokens, and it is text-only with thinking mode always on. The source reports benchmark results against Opus 4.8, Fable 5, GPT 5.6 Sol, and Qwen3.7-Max, and says the official Qwen3.8-Max API adds vision input and a 1M default context.

    Why it matters: The model card gives parameters, architecture, reasoning controls, and benchmark tables against named rival models, showing what an open release of this scale actually offers.

Aug 5

Aug 5Wed
  1. Prime Intellect BlogAI score75

    Prime Agent launches open-source self-improving RLM coding harness

    AIPrime Agent is a new open-source coding harness built on a persistent IPython kernel, a Recursive Language Model design, and Continual Harness state that the agent can create, read, update, and delete. Prime Intellect reports ARC-AGI-3 results of 95.5% RHAE Best@1 with Opus 5 and competitive long-context scores with the open-weights GLM-5.2 model.

    Why it matters: The post explains how the RLM and Continual Harness designs let an agent write code against its own context, sub-agents, and harness state, with benchmark evidence.

Aug 4

Aug 4Tue
  1. Zed BlogAI score65

    Zed Enables OS-Level Sandboxing by Default for Its Agent Panel

    AIZed's agent panel now sandboxes its terminal and fetch tools by default, starting in release 1.14, and the restrictions are enforced by the operating system rather than by agent instructions. By default the sandbox blocks writes outside project directories, writes to .git, and network requests, and agents can request temporary escalation with a stated reason. The post also notes that sandboxing covers only those tools and does not protect against other tools, external programs, or the regular built-in terminal.

    Why it matters: The post explains how OS-enforced sandboxing limits agent terminal and fetch access, and why fine-grained command rules fall short of it.

Aug 3

Aug 3Mon
  1. JetBrains AI BlogAI score52

    JetBrains Built a Central CLI to Control Spiraling AI Tool Costs

    AIJetBrains says its AI development expenses rose roughly 10x over six months as developers adopted three to five AI tools each. It built the JetBrains Central CLI, which routes third-party agent traffic through its AI platform so managers can set per-developer and team limits and view consumption reports. The CLI opened to early access on July 8 for anyone with JetBrains AI credits.

Aug 2

Aug 2Sun
  1. OpenRouter BlogAI score40

    OpenRouter Launches Ori Eval to Find the Best AI Model for Your App

    AIOpenRouter has released Ori Eval, an agent-driven tool that runs your app's prompts against candidate models and returns a comparison table of catch rate, latency, cost per PR, and pass/fail results. The tool asserts on called tools and grades open-ended answers with an LLM judge, pinning the harness and model during each run. Its evals are code files that can run in CI to block regressions and re-run when new models ship.

Jul 31

Jul 31Fri
  1. DeepSeek API NewsAI score67

    DeepSeek-V4-Flash API enters public beta with stronger agent benchmarks

    AIDeepSeek has released the DeepSeek-V4-Flash API in public beta, and developers can use the latest version by setting the model name to deepseek-v4-flash. The source reports agent benchmark results far above V4-Pro-Preview, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon verified. V4-Flash natively supports the Responses API format and is adapted for Codex, while V4-Pro and the APP/WEB models are unchanged.

    Why it matters: The release lists agent benchmark results against V4-Pro-Preview and notes Responses API support for Codex, which helps developers gauge the upgrade's practical effect on their workflows.

Jul 28

Jul 28Tue
  1. Augment Code BlogAI score39

    GPT-5.6 Sol Becomes Augment Cosmos's Default Model for Token Efficiency

    AIAugment Code has made GPT-5.6 Sol the default model in Cosmos, choosing it as the most token-efficient model to clear its pass-rate floor for long-horizon software engineering tasks. The company ranks models by cost per task rather than list price per million tokens, since retries on failed steps add token spend. Users can still select any model, and the default will change as more token-efficient models emerge.

  2. JetBrains AI BlogAI score60

    Ponytail Skill Cuts Claude Code Costs 10% But Not the Advertised 54%

    AIJetBrains tested the ponytail skill for Claude Code across 80 paired tasks and found a median 10.3% cost reduction, with p=0.004. Code written fell about 15% median versus the advertised 54%, reaching 31% on larger builds and little on already-lean tasks. No quality difference was detected, and the skill only self-activated when its ruleset was injected by a plugin hook.

    Why it matters: The benchmark separates advertised savings from measured results and shows the code cut depends on how much the baseline agent over-builds.