Skip to content

#Coding

Oct 8

TodayOct 8Thu9 items
  1. MarkTechPostAI score58

    JetBrains releases Mellum2.1, a 12B MoE open model for coding agents

    JetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.

  2. Leiphone (雷峰网)AI score62

    Claude Haiku 5.5 gains on computer use but still trails Sonnet 5.5 in terminal coding

    Anthropic released Claude Haiku 5.5, raising its OSWorld 2.1 score from 15.7% to 72.4% and supporting a 1 million token context window. The article notes Haiku 5.5 still scores 39.2% on Terminal-Bench 4.0 against Sonnet 5.5's 70.6%, and that prompts above 100,000 tokens are priced higher, so migration costs need to be measured on real workloads.

  3. OpenRouterAI score44

    Step 5 Preview from @StepFun_ai is live on OpenRouter. Their new flagship for agentic work: a sparse MoE (27B active, 600B total), 1M context, and text, image, and video input. Strong at coding and professional knowledge work, especially finance. Try it: https://openrouter.ai/stepfun/step-5-preview

    Step 5 Preview from @StepFun_ai is live on OpenRouter. Their new flagship for agentic work: a sparse MoE (27B active, 600B total), 1M context, and text, image, and video input. Strong at coding and professional knowledge work, especially finance. Try it: https://openrouter.ai/stepfun/step-5-preview

  4. JetBrains AI BlogAI score62

    JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning

    JetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.

    AIWhy it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.

Oct 7

Oct 7Wed
  1. IThome · AI (IT之家)AI score72

    Anthropic releases Claude Haiku 5.5, cutting run costs about 75% from Haiku 4.5

    Anthropic released Claude Haiku 5.5, which it calls the fastest, cheapest, and most capable Haiku model so far. On average it costs about 75% less to run than Haiku 4.5, with input at $0.10 and output at $0.50 per million tokens for requests up to 100,000 tokens. Anthropic also cut Sonnet 5.5's cache read price from $0.20 to $0.10 per million tokens, which it says lowers run costs by about 20% on many agent tasks.

  2. AWS Machine Learning BlogAI score56

    Claude Haiku 5.5 becomes available on Amazon Bedrock and Claude Platform on AWS

    Anthropic's Claude Haiku 5.5 is now available on Amazon Bedrock and Claude Platform on AWS. According to Anthropic, it is the fastest and most efficient model in the Claude 5.5 family and costs around 75 percent less than Claude Haiku 4.5 for most tasks. The post also covers pairing it with Claude Opus 5.5 as a subagent layer and provides Boto3, Converse, and Anthropic SDK examples for calling the model.

  3. Testing CatalogAI score62

    Anthropic releases Claude Haiku 5.5, its fastest and cheapest model

    Anthropic has released Claude Haiku 5.5, which the author describes as its fastest and cheapest model to date. The source says it costs about 75% less to run than Claude Haiku 4.5 and is the first Haiku model with an adjustable effort setting. The attached benchmark table reports Haiku 5.5 scores on tasks including computer use (OSWorld 2.1 offline subset, 72.4%) and Terminal-Bench 4.0 (39.2%), compared with Haiku 4.5 and other models.

  4. ClaudeAI score38

    Haiku 5.5 is designed for high-volume, cost-sensitive tasks. It reliably handles repetitive work like summaries and classification, and pairs well with Claude Opus 5.5 and Sonnet 5.5 as a subagent on coding work. It’s also fast enough for live customer support and browser use.

    Haiku 5.5 is designed for high-volume, cost-sensitive tasks. It reliably handles repetitive work like summaries and classification, and pairs well with Claude Opus 5.5 and Sonnet 5.5 as a subagent on coding work. It’s also fast enough for live customer support and browser use.

Oct 6

Oct 6Tue
  1. Guillaume LampleAI score42

    On important coding and agentic benchmarks such as DeepSWE, AutomationBench, and AA-Briefcase, ML4 matches the performance of the best open-weight models. It is SOTA on finance and legal workflows, as well as on complex multimodal grounding benchmarks. It can navigate complex terminal workflows, work across spreadsheets, slides, and PDFs, and reason over scientific and multimodal tasks. 4/n

    On important coding and agentic benchmarks such as DeepSWE, AutomationBench, and AA-Briefcase, ML4 matches the performance of the best open-weight models. It is SOTA on finance and legal workflows, as well as on complex multimodal grounding benchmarks. It can navigate complex terminal workflows, work across spreadsheets, slides, and PDFs, and reason over scientific and multimodal tasks. 4/n

  2. Latent SpaceAI score60

    Reflection launches Beam, a 501B-parameter open-weight coding model

    Reflection announced Beam, a text-only 501B-total, 23B-active MoE model for coding, agentic, and scientific work, trained from scratch with full weights under Apache 2.0 promised this month. Self-reported results include 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2, while the roundup notes that GLM 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash are generally ahead.

Oct 5

Oct 5Mon
  1. IThome · AI (IT之家)AI score49

    Reflection AI releases open-weight Beam model to rival DeepSeek and Kimi

    Reflection AI, an Nvidia-backed startup, released Beam, its first open-weight large model, aimed at coding and agent tasks. The company says Beam is comparable to Z.ai's GLM-5.2 and is approaching Qwen3.8-Max on coding and agent work. Beam has 501 billion total parameters, with 23 billion activated per task in a sparse architecture.

  2. NVIDIA AIAI score39

    The model behind this result is now on @huggingface 🤗 Nemotron-Labs-3-Competitive-Coding is a competitive-programming specialist model based on Nemotron-3-Ultra. https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4

    The model behind this result is now on @huggingface 🤗 Nemotron-Labs-3-Competitive-Coding is a competitive-programming specialist model based on Nemotron-3-Ultra. https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4

Oct 1

Oct 1Thu

Sep 30

Sep 30Wed
  1. indigoAI score81

    Google's Gemini 4 Argon debuts with limited access pending US government approval

    Google has announced Gemini 4 Argon, initially available only to trusted cyber defenders through its Fairwind Program while US government approval is pending. The author says the model is aimed at long-running software engineering, enterprise knowledge work, and cybersecurity tasks, with a 1 million token output limit. The post also gives promotional pricing of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward, alongside a benchmark comparison.

    AIWhy it matters: The post places Gemini 4 Argon's benchmark table beside GPT-6 Astra and Claude models, showing where each leads across coding, knowledge work, and cybersecurity tasks.

  2. Varun MohanAI score40

    Just announced Gemini 4 Argon, our new frontier model! It delivers frontier performance across the board on complex software tasks. Thousands of Googlers have already been using it in Antigravity internally and we’re excited by the feedback. It’s rolling out first to a set of trusted cyber defenders in our Fairwind Program with broader availability coming as soon as possible.

    Just announced Gemini 4 Argon, our new frontier model! It delivers frontier performance across the board on complex software tasks. Thousands of Googlers have already been using it in Antigravity internally and we’re excited by the feedback. It’s rolling out first to a set of trusted cyber defenders in our Fairwind Program with broader availability coming as soon as possible.

  3. Google AIAI score72

    Google announces Gemini 4 Argon, a frontier model with 1M output tokens

    Google AI announced Gemini 4 Argon, a new frontier model built for deep reasoning across long, complex workflows in software engineering, legal and finance knowledge work, and cybersecurity defense. Google says it is expanding the model's output token limit to 1M tokens. Argon is rolling out first to trusted cyber defenders in the Fairwind Program, with broader availability to follow as soon as possible.

    AIWhy it matters: The benchmark table compares Gemini 4 Argon against GPT-6 Astra and Claude models across knowledge work, coding, and multimodal tasks, showing where it leads and trails.

  4. Google DeepMindAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    Google DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    AIWhy it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  5. Google · Gemini appAI score91

    Google announces Gemini 4 Argon, rolling out first to trusted cyber defenders

    Google announced Gemini 4 Argon, a new frontier model rolling out first to trusted cyber defenders through its Fairwind Program. The model's output limit rises to 1M tokens from 64K, and its introductory API price is $2 per million input tokens and $10 per million output tokens. Google says broader availability to developers, enterprises, and consumers will follow after more testing of guardrails.

    AIWhy it matters: The post pairs benchmark claims with a phased access plan, pricing, and safety measures, which helps readers judge how quickly Argon may reach developers.

Sep 29

Sep 29Tue
  1. BAAI · new models on Hugging FaceAI score62

    BAAI releases AREX-2, a 27B agent model for self-improving long-horizon tasks

    BAAI released AREX-2, a 27B-parameter long-horizon agent model that improves solutions over multiple test-time rounds by proposing, measuring, reflecting, and revising. It was trained on machine-learning and algorithmic-programming tasks with verifiable feedback, and the source reports that this self-improvement transfers to deep research. The model is Apache License 2.0 licensed and has a 262,144-token context length.

    AIWhy it matters: The source compares AREX-2 against closed and open models on coding and deep-research benchmarks, showing how test-time self-improvement is measured across task types.

Sep 28

Sep 28Mon
  1. Cat WuAI score72

    Claude Sonnet 5.5 Lifts Claude Code Task Completion by About 30%

    Anthropic's Cat Wu says Claude Sonnet 5.5 lets Claude Code users complete about 30% more tasks than with Sonnet 5. The model needs fewer tokens for the same work, and in a leaf-raking tool-call demo it finished 24 seconds faster using 6K fewer tokens.

    AIWhy it matters: The post gives a measured Claude Code task-completion gain and a token-use example, showing what the model upgrade means for a coding agent workflow.

Sep 22

Sep 22Tue
  1. Mike KriegerAI score67

    Anthropic launches Claude Opus 5.5, leading in coding and knowledge work

    Anthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. According to the quoted launch post, it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. The author says it leads in coding and knowledge work and praises its writing quality.

    AIWhy it matters: The quoted launch post gives a concrete cost comparison, useful for weighing Opus 5.5 against earlier Opus and Fable 5.1 models for routine work.