Skip to content

Formats

Model releases Latest news

New models and updates: flagship releases, open weights, performance changes, and pricing changes.

183 picksPast 30 days: 78 itemsTotal: 1,043 items

Latest pick

Top picks archive · Page 2

Oct 2

Oct 2FriItems 21–40
  1. Hugging Face BlogAI score70

    Ai2 open-sources AstaBrief 8B, a fast model for generating cited research reports

    Ai2 released AstaBrief 8B, an open-weights model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. The model runs as Fast mode in Asta, averaging 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The post also describes filtering synthetic training data by citation density and building DPO pairs judged by two models that agreed.

    AIWhy it matters: The post explains how supervised fine-tuning, preference data, and citation-density filtering were used to build a cited-report model, which is useful for teams training their own models.

  2. Ai2 (Allen Institute for AI)AI score67

    Ai2 open-sources AstaBrief 8B, a fast open-weights scientific report model

    Ai2 released AstaBrief 8B, a model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. In Asta's Generate a report feature, Fast mode averages 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The model is built on Qwen3-8B with supervised fine-tuning and DPO, and institutions can run its open weights on their own infrastructure.

    AIWhy it matters: The post explains the data filtering and one-pass generation choices behind a fast open-weights report model, showing what worked and what did not.

Sep 30

Sep 30Wed
  1. indigoAI score81

    Google's Gemini 4 Argon debuts with limited access pending US government approval

    Google has announced Gemini 4 Argon, initially available only to trusted cyber defenders through its Fairwind Program while US government approval is pending. The author says the model is aimed at long-running software engineering, enterprise knowledge work, and cybersecurity tasks, with a 1 million token output limit. The post also gives promotional pricing of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward, alongside a benchmark comparison.

    AIWhy it matters: The post places Gemini 4 Argon's benchmark table beside GPT-6 Astra and Claude models, showing where each leads across coding, knowledge work, and cybersecurity tasks.

  2. Google AIAI score72

    Google announces Gemini 4 Argon, a frontier model with 1M output tokens

    Google AI announced Gemini 4 Argon, a new frontier model built for deep reasoning across long, complex workflows in software engineering, legal and finance knowledge work, and cybersecurity defense. Google says it is expanding the model's output token limit to 1M tokens. Argon is rolling out first to trusted cyber defenders in the Fairwind Program, with broader availability to follow as soon as possible.

    AIWhy it matters: The benchmark table compares Gemini 4 Argon against GPT-6 Astra and Claude models across knowledge work, coding, and multimodal tasks, showing where it leads and trails.

  3. Google DeepMindAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    Google DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    AIWhy it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  4. Google · Gemini appAI score91

    Google announces Gemini 4 Argon, rolling out first to trusted cyber defenders

    Google announced Gemini 4 Argon, a new frontier model rolling out first to trusted cyber defenders through its Fairwind Program. The model's output limit rises to 1M tokens from 64K, and its introductory API price is $2 per million input tokens and $10 per million output tokens. Google says broader availability to developers, enterprises, and consumers will follow after more testing of guardrails.

    AIWhy it matters: The post pairs benchmark claims with a phased access plan, pricing, and safety measures, which helps readers judge how quickly Argon may reach developers.

  5. Artificial Analysis ArticlesAI score75

    Gemini 4 Argon matches GPT-6 Astra on intelligence index at lower cost

    Artificial Analysis reports that Google's Gemini 4 Argon scores 53 on its Intelligence Index with high reasoning, matching GPT-6 Astra (max) and one point ahead of GPT-6.1 Sol (max). At the current 50% launch discount, its cost per task is $1.99, about 60% of GPT-6 Astra's $3.26, but the discount's end date is unconfirmed and standard pricing would raise it to $3.98. The model is being rolled out to selected users and is not publicly available.

    AIWhy it matters: The benchmark compares Gemini 4 Argon's cost per task and hallucination rate with GPT-6 Astra, showing where its value depends on a temporary 50% discount.

  6. Kling AI BlogAI score62

    Kling 4.0 enters early access with 30-second native video generation

    Kling 4.0 is entering early access, with a wider rollout planned for October, and Kling 4.0 Flash opens to Ultra Yearly subscribers on September 28. The update generates videos up to 30 seconds in a single pass, accepts up to 15 reference assets, and supports up to 10 keyframe images. Upcoming features include 10-bit HDR output at 4K and 1080p and video extension up to 2 minutes.

    AIWhy it matters: The post specifies concrete capability limits such as 30-second native generation, up to 15 references, and 10 keyframes, which help users judge fit for production workflows.

Sep 29

Sep 29Tue
  1. Tibor BlahoAI score78

    OpenAI's DevDay 2026 brings dots agents, GPT-6.1 Sol, and Ultrafast speed tier

    OpenAI announced more than 20 updates at DevDay 2026, including dots always-on agents, GPT-6.1 Sol, Ultrafast token generation, ChatGPT Space, and a $500/month Pro 500 plan. GPT-6.1 Sol is priced at $2 input and $10 output per 1M tokens and is available in the API as gpt-6.1-sol. Ultrafast generates tokens up to 8x faster in Codex and up to 6x faster in the API.

    AIWhy it matters: The post lists dozens of OpenAI DevDay 2026 changes across models, agents, plans, and APIs, useful for scanning what shipped and who gets access.

  2. BAAI · new models on Hugging FaceAI score62

    BAAI releases AREX-2, a 27B agent model for self-improving long-horizon tasks

    BAAI released AREX-2, a 27B-parameter long-horizon agent model that improves solutions over multiple test-time rounds by proposing, measuring, reflecting, and revising. It was trained on machine-learning and algorithmic-programming tasks with verifiable feedback, and the source reports that this self-improvement transfers to deep research. The model is Apache License 2.0 licensed and has a 262,144-token context length.

    AIWhy it matters: The source compares AREX-2 against closed and open models on coding and deep-research benchmarks, showing how test-time self-improvement is measured across task types.

  3. Artificial Analysis ArticlesAI score78

    GPT-6.1 Sol replaces GPT-6 Sol with near-Astra intelligence at lower cost

    Artificial Analysis reports that GPT-6.1 Sol replaces GPT-6 Sol after seven days and scores 1 point below GPT-6 Astra on the Intelligence Index. At max effort it costs $0.72 per Intelligence Index task, compared with $3.26 for GPT-6 Astra and $1.05 for GPT-6 Sol. Its pricing matches GPT-6 Sol at $2/$10 per million input/output tokens, but it uses about 10-30% more output tokens.

    AIWhy it matters: The source compares GPT-6.1 Sol against GPT-6 Sol, GPT-5.6 Sol, and GPT-6 Astra on cost per task and token use, helping readers weigh performance against price.

Sep 28

Sep 28Mon
  1. Mike KriegerAI score67

    Anthropic releases Claude Sonnet 5.5, 30% faster and up to 30% cheaper than Sonnet 5

    Anthropic has released Claude Sonnet 5.5, the second model in the Claude 5.5 family. The company says it is more than 30% faster than Sonnet 5 and costs up to 30% less for most work.

    AIWhy it matters: The post gives concrete speed and price changes against Sonnet 5, which helps readers judge whether the upgrade fits their workloads and budgets.

  2. Cat WuAI score72

    Claude Sonnet 5.5 Lifts Claude Code Task Completion by About 30%

    Anthropic's Cat Wu says Claude Sonnet 5.5 lets Claude Code users complete about 30% more tasks than with Sonnet 5. The model needs fewer tokens for the same work, and in a leaf-raking tool-call demo it finished 24 seconds faster using 6K fewer tokens.

    AIWhy it matters: The post gives a measured Claude Code task-completion gain and a token-use example, showing what the model upgrade means for a coding agent workflow.

Sep 27

Sep 27Sun
  1. Claude Apps Release NotesAI score65

    Anthropic launches Claude Sonnet 5.5 as second Claude 5.5 model

    Anthropic has launched Claude Sonnet 5.5, the second model in its Claude 5.5 family. The company describes it as a faster, lower-cost complement to Claude Opus 5.5, and points readers to a blog post for more information.

    AIWhy it matters: The release note places Sonnet 5.5 beside Opus 5.5 in the Claude 5.5 family, clarifying which model suits speed and cost needs.

  2. Amp NewsAI score67

    Amp switches its default medium mode to Claude Opus 5.5

    Amp now uses Claude Opus 5.5 for its medium mode by default, replacing GPT-5.6 Sol, while ChatGPT subscribers can keep medium pinned to GPT-5.6 Sol. In Amp's internal evals, Opus 5.5 solved 65% of tasks versus 61% for GPT-5.6 Sol and 56% for Opus 5, at lower cost, and it runs at high reasoning effort because xhigh and max cost more without scoring better.

    AIWhy it matters: The source reports internal eval scores, cost comparisons, and usage guidance for choosing reasoning effort, helping developers decide which model and setting to run.

  3. Tibor BlahoAI score85

    OpenAI releases GPT-6 Sol and Luna as Anthropic launches Claude Opus 5.5

    OpenAI released GPT-6 Sol and Luna, priced 50 percent below GPT-5.6 promo API pricing, and rolling out in ChatGPT Work, Codex and the API, not yet in regular Chat. Anthropic released Claude Opus 5.5, described as roughly Claude Fable 5.1 level for 40 percent less than Opus 5 and over 30 percent faster, with Sonnet 5.5 and Haiku 5.5 due in coming weeks.

    AIWhy it matters: The recap puts OpenAI and Anthropic releases side by side, with pricing and capability claims that help compare the two launches.

  4. Xiaomi MiMoAI score62

    Xiaomi MiMo Explains Fixing Tool-Call Repetition in MiMo-V2.6 Models

    Xiaomi MiMo reports that tool-call repetition in MiMo-V2.6 reached over 0.05% of responses across agent harnesses, causing stalled agents and wasted context. The team traced the cause to an RL flooding penalty set at 32 calls per turn, which missed smaller excess behavior, and replaced the approach with a specialized teacher distilled via MOPD. Repetition rates for both Pro and Flash dropped substantially, at roughly $90,000 versus an estimated $2.31 million for the alternative fix.

    AIWhy it matters: The post traces an agent failure to a reward blind spot and compares the costs of two fixes, offering a transferable debugging method for RL-trained tool-calling models.

Sep 25

Sep 25Fri
  1. Anthropic ResearchAI score67

    Claude computes a nine-loop physics amplitude that experts had not reached

    Anthropic researchers used Claude Science to compute the nine-loop six-particle amplitude in planar N=4 super Yang-Mills, a toy-model result that physicist Lance Dixon checked. The work reportedly cost roughly one or two thousand dollars, with about $100 of compute for the bootstrap calculation, and a similar result was reached by Song He's group.

    AIWhy it matters: The guest post shows a frontier physics calculation done with modest compute, which helps readers gauge what current AI can handle in research and what it still cannot.

Sep 24

Sep 24Thu
  1. Google · Gemini appAI score62

    Google launches Gemini 3.8 Live with Live Avatar for enterprises

    Google introduced Gemini 3.8 Live with Live Avatar, which adds a visual persona with lip-syncing and expressions to its live dialogue models. The feature is available in Gemini Enterprise and supports 97 languages, with custom avatars available through enterprise allowlisting. Google says all output is watermarked with SynthID.

    AIWhy it matters: The post specifies enterprise availability, custom avatar allowlisting, and 97-language support, which clarifies who can use the feature and how far it reaches.

Sep 23

Sep 23Wed
  1. Google Developers BlogAI score62

    Google reproduces Olmo 3 7B pre-training in MaxText on TPUs

    Google Developers reproduced Ai2's Olmo 3 7B from scratch in MaxText on Google Cloud TPUs, covering both the stage-1 pre-training run and the stage-2 mid-training anneal. The match was checked on held-out C4 loss, an 8-task accuracy suite, multi-domain perplexity, and token-level KL, not just the training loss curve. The post also describes a data-loader bug that made training loss look better than the reference while held-out metrics did not move.

    AIWhy it matters: The post documents how a faithful reproduction was verified on held-out metrics, including a data bug that training loss alone would have hidden.