Skip to content

Formats · Latest news

Model releases

New models and updates: flagship releases, open weights, performance changes, and pricing changes.

281 top picks all-time · 125 in the past 30 days · chosen from 1,166 items collected all-time

Latest pick

Top picks archive · Page 13

Top picks 241–260 of 281

Apr 14

Apr 14Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score78

    Moonshot AI releases open-source Kimi K2.6 multimodal agentic model

    AIMoonshot AI released Kimi K2.6, an open-source native multimodal agentic model with 1T total and 32B activated parameters and a 256K context length. The model card reports benchmark results against GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro across agentic, coding, reasoning, and vision tasks, and supports swarms of up to 300 sub-agents.

    Why it matters: The model card gives specific agent swarm scale, context length, and benchmark comparisons against several frontier models, useful for judging its coding and agent capabilities.

Apr 8

Apr 8Wed
  1. MiniMax · new models on Hugging FaceOfficialAI score78

    MiniMax releases open-weight MiniMax-M2.7 with agent and coding gains

    AIMiniMax has released MiniMax-M2.7 on Hugging Face, describing it as its first model to participate in its own evolution. The source reports 56.22% on SWE-Pro, 46.3% on Toolathon, and 62.7% on MM ClawBench, and says an internal version autonomously optimized a programming scaffold over 100+ rounds for a 30% performance improvement.

    Why it matters: The source ties its benchmark claims to a self-evolution process and a named comparison set, which helps readers weigh how the reported gains were achieved.

Apr 7

Apr 7Tue
  1. Dario AmodeiXAI score62

    Anthropic's Dario Amodei says new Mythos Preview model shows a large jump in cyber capabilities

    AIDario Amodei says the company has tracked growing cyber capabilities in AI models for years, which arise from their general coding proficiency. He states that the new model, Mythos Preview, represents a particularly large step up in those capabilities.

    Why it matters: The post links rising cyber capability to general coding skill, and names a notable jump in a new model, Mythos Preview.

Apr 3

Apr 3Fri
  1. Z.ai (GLM) · new models on Hugging FaceOfficialAI score73

    Z.ai releases GLM-5.1, a flagship model for agentic engineering

    AIZ.ai has released GLM-5.1, its next-generation flagship model for agentic engineering, with stronger coding than GLM-5. The model is described as staying effective over longer agentic tasks, sustaining optimization over hundreds of rounds and thousands of tool calls. The release lists benchmark results including SWE-Bench Pro at 58.4 and Terminal-Bench 2.0 at 63.5, and local deployment is supported through SGLang, vLLM, xLLM, Transformers, and KTransformers.

    Why it matters: The release gives benchmark tables against several rival models, letting readers compare GLM-5.1's coding and agentic results with GLM-5 and frontier systems.

Mar 31

Mar 31Tue
  1. Mistral AI · new models on Hugging FaceOfficialAI score76

    Mistral Medium 3.5 releases as a 128B dense merged model with vision

    AIMistral AI released Mistral Medium 3.5, a dense 128B model with a 256k context window that handles instruction-following, reasoning, and coding in a single set of weights. It replaces Mistral Medium 3.1, Magistral, and Devstral 2, and reasoning effort is configurable per request. The model accepts text and image input and is released under a Modified MIT License that excludes companies with large revenue.

    Why it matters: The release merges instruction, reasoning, and coding into one 128B model with per-request reasoning control, giving developers one set of weights to compare against separate specialized models.

Mar 26

Mar 26Thu
  1. Guillaume Lample @ NeurIPS 2024XAI score62

    Mistral releases Voxtral TTS text-to-speech model with open weights

    AIMistral has released Voxtral TTS, a text-to-speech model, alongside a blog post, a playground, a technical report, and model weights on Hugging Face. The post itself contains only links and no further details about the model's capabilities.

    Why it matters: The post links a playground, technical report, and open model weights, letting readers test and verify the release themselves.

  2. Guillaume Lample @ NeurIPS 2024XAI score62

    Mistral releases Voxtral TTS, its first open-weight speech model

    AIMistral's Voxtral TTS is its first speech model, presented as an open-weight text-to-speech model that reportedly delivers SOTA performance at significantly lower cost with very low latency. It combines autoregressive generation of semantic speech tokens with flow-matching for acoustic tokens, and a technical report on its training methodology is being released.

    Why it matters: The post names Voxtral TTS's architecture and a technical report, giving readers a concrete basis for comparing its speech generation method with other text-to-speech systems.

    Image from @GuillaumeLample's post

Mar 17

Mar 17Tue
  1. Xiaomi MiMoOfficialAI score71

    Xiaomi releases MiMo-V2-Omni, an omni-modal model for agentic tasks

    AIXiaomi introduces MiMo-V2-Omni, a single model that fuses image, video, and audio encoders into a shared backbone with native tool calling and UI grounding. The company reports benchmark results against Gemini 3 Pro, Claude Opus 4.6, and GPT 5.2, and demonstrates browser-based shopping and video-publishing workflows run through the OpenClaw agent scaffold. It also states the model supports over 10 hours of continuous audio understanding.

    Why it matters: The page gives benchmark comparisons, a driving-risk demo, and browser-task walkthroughs, letting readers check how far the omni-modal claims extend into agent use.

  2. MiniMax BlogOfficialAI score63

    MiniMax M2.7 takes part in its own model and harness evolution

    AIMiniMax says M2.7 is its first model to deeply participate in its own evolution, building agent harnesses and running reinforcement learning experiment workflows. The post reports 56.22% on SWE-Pro, 55.6% on VIBE-Pro, 57.0% on Terminal Bench 2, and a 30% improvement on an internal evaluation set after more than 100 autonomous optimization rounds. It also states that M2.7 handles 30%-50% of its research team's workflow, though human researchers still make critical decisions.

    Why it matters: The post ties M2.7's self-evolution claims to specific benchmark numbers and workflow details, helping readers judge how much of the iteration loop is autonomous.

  3. Xiaomi MiMoOfficialAI score80

    Xiaomi MiMo-V2-Pro Flagship Model Targets Agent Workloads With 1M Context

    AIXiaomi announced MiMo-V2-Pro, a flagship foundation model for agent workloads with over 1T total parameters, 42B active, and up to 1M-token context. It ranks 8th worldwide and 2nd among Chinese LLMs on the Artificial Analysis Intelligence Index, and its API is publicly available with usage-tiered pricing.

    Why it matters: The post gives benchmark placements, parameter scale, context length, and tiered API pricing, so readers can compare it against Claude and GPT models on concrete terms.

  4. Xiaomi MiMoOfficialAI score68

    Xiaomi releases MiMo-V2-TTS, a speech model with controllable emotion and singing

    AIXiaomi has launched MiMo-V2-TTS, a speech synthesis model that lets users describe the desired voice style in plain language. The model also supports dialects, character voices, non-verbal sounds such as coughs and sighs, and singing within one model. It was pretrained on over 100 million hours of speech data and refined with multi-dimensional reinforcement learning.

    Why it matters: The source gives concrete controls for emotion, dialect, singing, and non-verbal sounds, showing how a voice model can be directed through plain-language style prompts.

Mar 11

Mar 11Wed
  1. Mistral AI · new models on Hugging FaceOfficialAI score62

    Mistral AI releases Leanstral-2603, an open-source Lean 4 proof agent

    AIMistral AI released Leanstral 119B A6B on Hugging Face as an open-source code agent for Lean 4 proof engineering. The model uses 128 experts with 4 active per token, 6.5B activated parameters, a 256k token context window, and accepts text and image input under the Apache 2.0 license. The page also documents vLLM server deployment and Mistral Vibe integration.

    Why it matters: The source specifies Leanstral's 119B MoE architecture, 256k context, Apache 2.0 license, and vLLM setup, showing how the Lean 4 proof agent could be deployed locally.

Mar 5

Mar 5Thu
  1. Nick TurleyXAI score62

    GPT-5.4 Thinking rolls out to ChatGPT with mid-response interrupts

    AIGPT-5.4 Thinking is rolling out to ChatGPT, and users can now interrupt it before it produces the final answer. Users can steer the response while it is still working rather than sending multiple follow-up turns. The update also improves deep web research and long-context reasoning, which the post says helps specific questions arrive faster and stay focused.

    Why it matters: The post names the new interrupt control and the research and long-context gains, showing how this change affects steering responses in ChatGPT.

Mar 4

Mar 4Wed
  1. Mistral AI · new models on Hugging FaceOfficialAI score67

    Mistral Small 4 unifies instruct, reasoning, and coding in one open model

    AIMistral Small 4 combines instruct, reasoning, and Devstral capabilities in one multimodal model with 119B total parameters, 6.5B active per token, and a 256k context window. The source reports a 40% reduction in latency-optimized end-to-end completion time and 3x more requests per second in throughput-optimized setups versus Mistral Small 3. It is released under Apache 2.0 and supports reasoning mode toggling per request.

    Why it matters: The source lists architecture, context length, and mode-switching controls, letting readers compare this release's design with earlier Mistral Small models.

Mar 3

Mar 3Tue
  1. Nick TurleyXAI score62

    OpenAI rolls out GPT-5.3 Instant in ChatGPT with fewer refusals and disclaimers

    AIOpenAI's Nick Turley announced that GPT-5.3 Instant is rolling out in ChatGPT starting today. The update responds to feedback that GPT-5.2 was sometimes too cautious, over-caveated, and less natural in conversation, with fewer unnecessary refusals, fewer defensive disclaimers, and more direct answers.

    Why it matters: The post names the specific complaints about GPT-5.2 and the behavior changes made in response, which shows how user feedback shaped this update.

Feb 26

Feb 26Thu
  1. Oriol VinyalsXAI score60

    Nano Banana 2 debuts at #1 in Image Arena text-to-image ranking

    AINano Banana 2, officially released as Gemini 3.1 Flash Image Preview, ranks first in Image Arena text-to-image with a score of 1279. The quoted post says it also ties for first in single-image editing at 1407 and costs $0.067 per image, about half the price of Nano Banana Pro.

    Why it matters: The quoted leaderboard figures and per-image price give concrete reference points for comparing this image model against Nano Banana Pro and GPT-Image-1.5.

    Image from @OriolVinyalsML's post
  2. Nano Banana 2.1OfficialAI score67

    Google introduces Nano Banana 2, its best image generation and editing model

    AINano Banana announces Nano Banana 2, which it describes as its best image generation and editing model yet. The model can be tried in the Gemini app, Google AI Studio, and other places the post does not specify.

    Why it matters: The post names the access points for Nano Banana 2, which helps readers see where the image generation and editing model can be tried.

Feb 19

Feb 19Thu
  1. Yi TayXAI score78

    Google releases Gemini 3.1 Pro, reporting 77.1% on ARC-AGI-2

    AIGoogle has released Gemini 3.1 Pro, reporting 77.1% on ARC-AGI-2 and more than twice the score of Gemini 3 Pro on that benchmark. The model is rolling out to developers in preview through the Gemini API and Google AI Studio, to enterprises via Vertex AI and Gemini Enterprise, and to consumers in the Gemini app and NotebookLM.

    Why it matters: The post pairs the release with a benchmark table comparing Gemini 3.1 Pro against Gemini 3 Pro, Claude Sonnet 4.6, Claude Opus 4.6, and GPT-5.2 on reasoning and coding tasks.

Feb 17

Feb 17Tue
  1. Eugene YanXAI score72

    Claude Sonnet 4.6 released with upgrades and 1M token context window

    AIAnthropic's Claude Sonnet 4.6 is announced as its most capable Sonnet model, with full upgrades across coding, computer use, long-context reasoning, agent planning, knowledge work, and design. It also features a 1M token context window in beta. The author notes that the model is versatile across classification, coding, computer use, and autonomous agents by adjusting effort and thinking modes.

    Why it matters: The post places Sonnet 4.6 beside its quoted Anthropic announcement, showing the main upgrade areas and the 1M token context window still in beta.

Feb 12

Feb 12Thu
  1. MiniMax · new models on Hugging FaceOfficialAI score88

    MiniMax releases M2.5 model with 80.2% on SWE-Bench Verified

    AIMiniMax has released M2.5, which it says reaches 80.2% on SWE-Bench Verified and 76.3% on BrowseComp with context management. The company reports 37% faster end-to-end runtime than M2.1 on SWE-Bench Verified and prices M2.5 at $1 per hour at 100 tokens per second, with a 50 tokens per second version at $0.30 per hour. Weights are available on Hugging Face, with inference support listed for SGLang, vLLM, Transformers, and KTransformers.

    Why it matters: The source gives benchmark scores against Claude and GPT models plus per-task token and runtime figures, so readers can weigh the cost-speed tradeoff directly.