Skip to contentSkip to stories
Updated

#Coding

Oct 8

  1. JetBrains AI BlogAI score62

    JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning

    AIJetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.

    Why it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.

Sep 30

  1. indigoAI score81

    Google's Gemini 4 Argon debuts with limited access pending US government approval

    AIGoogle has announced Gemini 4 Argon, initially available only to trusted cyber defenders through its Fairwind Program while US government approval is pending. The author says the model is aimed at long-running software engineering, enterprise knowledge work, and cybersecurity tasks, with a 1 million token output limit. The post also gives promotional pricing of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward, alongside a benchmark comparison.

    Why it matters: The post places Gemini 4 Argon's benchmark table beside GPT-6 Astra and Claude models, showing where each leads across coding, knowledge work, and cybersecurity tasks.

  2. Google AIAI score72

    Google announces Gemini 4 Argon, a frontier model with 1M output tokens

    AIGoogle AI announced Gemini 4 Argon, a new frontier model built for deep reasoning across long, complex workflows in software engineering, legal and finance knowledge work, and cybersecurity defense. Google says it is expanding the model's output token limit to 1M tokens. Argon is rolling out first to trusted cyber defenders in the Fairwind Program, with broader availability to follow as soon as possible.

    Why it matters: The benchmark table compares Gemini 4 Argon against GPT-6 Astra and Claude models across knowledge work, coding, and multimodal tasks, showing where it leads and trails.

  3. Google DeepMindAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    AIGoogle DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    Why it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  4. Google · Gemini appAI score91

    Google announces Gemini 4 Argon, rolling out first to trusted cyber defenders

    AIGoogle announced Gemini 4 Argon, a new frontier model rolling out first to trusted cyber defenders through its Fairwind Program. The model's output limit rises to 1M tokens from 64K, and its introductory API price is $2 per million input tokens and $10 per million output tokens. Google says broader availability to developers, enterprises, and consumers will follow after more testing of guardrails.

    Why it matters: The post pairs benchmark claims with a phased access plan, pricing, and safety measures, which helps readers judge how quickly Argon may reach developers.

Sep 29

  1. BAAI · new models on Hugging FaceAI score62

    BAAI releases AREX-2, a 27B agent model for self-improving long-horizon tasks

    AIBAAI released AREX-2, a 27B-parameter long-horizon agent model that improves solutions over multiple test-time rounds by proposing, measuring, reflecting, and revising. It was trained on machine-learning and algorithmic-programming tasks with verifiable feedback, and the source reports that this self-improvement transfers to deep research. The model is Apache License 2.0 licensed and has a 262,144-token context length.

    Why it matters: The source compares AREX-2 against closed and open models on coding and deep-research benchmarks, showing how test-time self-improvement is measured across task types.

Sep 28

  1. Cat WuAI score72

    Claude Sonnet 5.5 Lifts Claude Code Task Completion by About 30%

    AIAnthropic's Cat Wu says Claude Sonnet 5.5 lets Claude Code users complete about 30% more tasks than with Sonnet 5. The model needs fewer tokens for the same work, and in a leaf-raking tool-call demo it finished 24 seconds faster using 6K fewer tokens.

    Why it matters: The post gives a measured Claude Code task-completion gain and a token-use example, showing what the model upgrade means for a coding agent workflow.

Sep 22

  1. Mike KriegerAI score67

    Anthropic launches Claude Opus 5.5, leading in coding and knowledge work

    AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. According to the quoted launch post, it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. The author says it leads in coding and knowledge work and praises its writing quality.

    Why it matters: The quoted launch post gives a concrete cost comparison, useful for weighing Opus 5.5 against earlier Opus and Fable 5.1 models for routine work.

Sep 21

  1. Xiaomi MiMoAI score78

    Xiaomi releases open-weight MiMo-V2.6 Pro and Flash omnimodal models

    AIXiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.

    Why it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

Sep 20

  1. xAI News (Grok)AI score72

    xAI releases Grok 4.7, its most capable model for coding and knowledge work

    AIxAI released Grok 4.7, which it calls its most capable model for coding and knowledge work, built on a larger base model than Grok 4.6 and trained with a longer reinforcement learning run. It is priced from $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, and is available in Cursor, Grok Build, and the Grok API. xAI reports gains on CursorBench 4.0 (46.3%) and AA Briefcase v1.1 (1,657) over Grok 4.6, and says it posts the strongest safety results it has tested on refusals and jailbreak resistance.

    Why it matters: The release pairs a new base model with benchmark tables against named rivals and pricing, letting readers compare its coding and office-work gains against Grok 4.6 and frontier models.

Sep 11

  1. Baseten BlogAI score62

    DeepSeek-V4.1-Flash arrives on Baseten with a split prefill architecture

    AIDeepSeek released open weights for V4.1-Flash, which Baseten now offers through its Model APIs. The model has 552B total parameters, 8B active for prefill and 16B for decode, a 1M token context window, and text plus image input. Its Causal Encoder-Decoder design runs only the encoder during prefill and reuses a projected KV cache, and the source reports the global KV cache at a quarter of V4-Flash's memory.

    Why it matters: The post explains how the CED architecture splits prefill and decode compute and cuts KV cache memory, which matters for coding agent costs.

  2. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score72

    Shanghai AI Lab releases Atria Dawn Preview, an agentic model built on GLM-5.2

    AIShanghai Artificial Intelligence Laboratory has released Atria Dawn Preview, an agentic model built on the 744B-parameter MoE GLM-5.2 foundation model, with a 256K context window. The release page reports benchmark results across search, coding, tool use, productivity, and cybersecurity, and describes text-only setup for Codex and Claude Code.

    Why it matters: The release page gives a full benchmark table against named rivals and setup steps for Codex and Claude Code, useful for anyone evaluating agentic models.

Sep 10

  1. Cognition Blog (Devin, Windsurf)AI score66

    Cognition releases SWE-2, a coding model trained with cost-penalized RL

    AICognition introduces SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while costing 64% less. The post attributes the gains to an RL algorithm that trains all reasoning-effort levels in one run, with cost penalties tuned to the base model's Pareto frontier. SWE-2 is available starting today in Devin Desktop and CLI, with rollout to Devin Web and Fusion.

    Why it matters: The post explains how the cost penalty and length-weighted baseline are derived, which helps readers judge the tradeoffs in coding model post-training.

Sep 6

  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score62

    OpenBMB releases MiniCPM5-2B, a 2B open-source model with open training data

    AIOpenBMB has released MiniCPM5-2B, a dense 2B Transformer built for on-device and resource-constrained deployment, with an average score of 53.9 in its comparison set. The release also opens the UltraData datasets behind it, including UltraX, UltraData-Code, UltraData-SFT-Agent-2609 and UltraData-RL-2609, and includes GGUF, MLX, GPTQ and DSpark variants for common runtimes.

    Why it matters: The release pairs a 2B model with open training datasets and reports per-benchmark comparisons against named same-size and larger models, letting readers check the claims directly.

Sep 3

  1. Mark ChenAI score80

    Mark Chen announces GPT-6 Astra with computer use and agent oversight

    AIOpenAI researcher Mark Chen announced GPT-6 Astra, which he described as the company's most capable and aligned model yet. He said it can build and test software, work across apps on a computer, and help with open scientific problems. The post also highlights improved computer use compared with Operator and stronger monitoring that can stop potentially unauthorized agent actions.

    Why it matters: The post links a named model release to specific capabilities like computer use and aligned agent behavior, giving readers concrete claims to check against the model.

Sep 2

  1. Google AI StudioAI score78

    Google releases Gemini 3.8 Flash and restricted 3.8 Flash Cyber model

    AIGoogle introduces Gemini 3.8 Flash for coding, agentic tasks, and multi-step reasoning, priced at $0.75 per million input tokens and $3.75 per million output tokens during the introductory period. Gemini 3.8 Flash Cyber targets vulnerability detection and automated patching and is available only to trusted defenders through the new Fairwind Program. The introductory price expires December 31, 2026, after which $1.50 and $7.50 per million tokens apply.

    Why it matters: The post separates a general coding and agent model from a restricted cyber variant, showing how one shared core is deployed under different access and safety tiers.

Sep 1

  1. Anthropic · YouTubeAI score78

    Anthropic releases Claude Fable 5.1, an upgrade to its most capable model class

    AIAnthropic has released Claude Fable 5.1, the latest upgrade to its most capable class of models, and it is available everywhere today. The company says it handles complex, long-running, multi-step work and avoids shortcuts when fixing root causes of software issues. At lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost, according to Anthropic's benchmarks.

    Why it matters: The source names the upgraded model class and its cost tradeoff at lower effort levels, which helps readers weigh it against the earlier version for their own workloads.

Aug 31

  1. Claude Apps Release NotesAI score72

    Anthropic launches Claude Fable 5.1 and Claude Mythos 5.1 models

    AIAnthropic has launched Claude Fable 5.1 and Claude Mythos 5.1, which it describes as the world's most advanced models for coding and knowledge work. The release notes link to a blog post with more details, but the notes themselves give no benchmarks or specifications.

    Why it matters: The source names two new model versions and points to a companion blog post, so readers can compare the release details there.

Aug 27

  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score65

    OpenBMB releases MiniCPM5-2B-SFT, a 2B open model with SFT-only checkpoint

    AIOpenBMB released MiniCPM5-2B-SFT, an SFT-only BF16 checkpoint taken before RL and OPD, within its MiniCPM5-2B series. The model is a 2B dense Transformer built for on-device and local deployment, with 131,072-token context and the same training recipe as the final release.

    Why it matters: The source gives concrete benchmark averages against same-size and larger models, plus released training data and multiple deployment formats, useful for judging a compact on-device model.