Skip to content

Formats · Latest news

Model releases

New models and updates: flagship releases, open weights, performance changes, and pricing changes.

281 top picks all-time · 125 in the past 30 days · chosen from 1,171 items collected all-time

Latest pick

Top picks archive · Page 14

Top picks 261–280 of 281

Feb 10

Feb 10Tue
  1. Z.ai (GLM) · new models on Hugging FaceOfficialAI score72

    Z.ai releases GLM-5, a 744B-parameter open model for agentic engineering

    AIZ.ai launches GLM-5, scaling from 355B to 744B total parameters with 40B active and pre-training data from 23T to 28.5T tokens. The model integrates DeepSeek Sparse Attention to reduce deployment cost and reports strong results on reasoning, coding, and agentic benchmarks against GLM-4.7, DeepSeek-V3.2, Kimi K2.5, and several frontier models.

    Why it matters: The source gives concrete scale, data, and benchmark comparisons against named frontier models, showing where GLM-5 sits among open-source and proprietary systems.

Feb 4

Feb 4Wed
  1. Guillaume Lample @ NeurIPS 2024XAI score62

    Mistral releases Mini Transcribe 2 and realtime transcription with open weights

    AIMistral announces Mini Transcribe 2 via API at $0.003 per minute and a realtime transcription option at $0.006 per minute. The realtime model's open weights are published on Hugging Face, alongside a realtime demo and a blog post.

    Why it matters: The post gives per-minute API prices and open weights for a realtime tier, letting developers compare transcription costs against existing speech-to-text options.

  2. Guillaume Lample @ NeurIPS 2024XAI score62

    Mistral's Voxtral Realtime streams speech with sub-200ms latency and open weights

    AIVoxtral Realtime is a natively streaming speech model for voice agents and live applications, with latency configurable down to sub-200ms. At 480ms it stays within 1-2% WER of the offline model, and the weights are released under Apache 2.0. The attached FLEURS chart compares word error rates across latency settings for ten languages, including Chinese.

    Why it matters: The post gives latency and accuracy tradeoff figures for a streaming speech model, helping readers judge whether it suits real-time voice agents.

    Image from @GuillaumeLample's post
  3. Guillaume Lample @ NeurIPS 2024XAI score62

    Mistral releases Voxtral 2 transcription models with real-time option

    AIMistral announces Voxtral 2 with two transcription models: Voxtral Realtime, released under an Apache 2 license with latency configurable to sub-200 ms, and Voxtral Mini Transcribe 2, which adds speaker diarization, word-level timestamps, and context biasing. The models support 13 languages and are available through the Mistral API, which the post describes as one of the most cost-effective transcription APIs on the market. The attached chart shows word error rates on FLEURS across Italian, Spanish, English, German, Portuguese, French, Russian, Dutch, and Chinese at several latency settings.

    Why it matters: The post names two transcription models, their Apache 2 license, and the sub-200 ms latency option, showing what changes for real-time speech workflows.

    Image from @GuillaumeLample's post
  4. Intern Large ModelsOfficialAI score60

    Intern-S1-Pro: 1T MoE open-source multimodal scientific reasoning model released

    AIIntern Large Models introduces Intern-S1-Pro, a 1T-parameter MoE open-source multimodal scientific reasoning model with 1T-A22B active configuration. The post claims competitive scientific reasoning against leading closed-source models and supports vLLM and SGLang, with weights on Hugging Face and code on GitHub.

    Why it matters: The post pairs a 1T MoE open-source scientific reasoning model with benchmark tables against closed models, letting readers compare claimed strengths directly.

    Image from @intern_lm's post

Jan 29

Jan 29Thu
  1. Z.ai (GLM) · new models on Hugging FaceOfficialAI score60

    Z.ai releases open-source GLM-OCR multimodal document model

    AIZ.ai has released GLM-OCR, a 0.9B-parameter multimodal OCR model for complex document understanding, under the MIT License. The model scores 94.62 on OmniDocBench V1.5 and supports deployment through vLLM, SGLang, and Ollama, with an official SDK for document parsing.

    Why it matters: The page gives benchmark scores, a 0.9B parameter size, and supported serving frameworks, which help readers weigh OCR deployment options against heavier alternatives.

Jan 23

Jan 23Fri
  1. Mistral AI · new models on Hugging FaceOfficialAI score67

    Mistral Small 4 unifies instruct, reasoning, and coding in one open model

    AIMistral Small 4 is a 119B-parameter MoE model with 6.5B active per token and a 256k context window, combining instruct, reasoning, and Devstral-style coding in one model. It accepts text and image input, lets users set reasoning_effort per request, and is released under Apache 2.0. The model card reports a 40% latency reduction and 3x throughput versus Mistral Small 3 in its tested setups, and its benchmark chart shows reasoning scores on GPQA Diamond, MMLU Pro, AIME-style text tasks, and MMMU-Pro.

    Why it matters: The model card names concrete architecture, context, and licensing details, letting readers compare its reasoning toggle and efficiency claims against other open models.

Jan 21

Jan 21Wed
  1. Mistral AI · new models on Hugging FaceOfficialAI score65

    Mistral releases open-weight Voxtral Mini 4B Realtime 2602 speech model

    AIMistral AI released Voxtral Mini 4B Realtime 2602, a multilingual realtime speech-transcription model with 13 supported languages under the Apache 2.0 license. The model has a configurable transcription delay from 240ms to 2.4s, and it matches leading offline open-source models at a 480ms delay. The source says it is optimized for on-device deployment and is currently supported only in vLLM.

    Why it matters: The source specifies the 480ms delay operating point, 4B size, Apache 2.0 license, and vLLM serving path, which matter for teams weighing realtime transcription deployment.

Jan 19

Jan 19Mon
  1. Z.ai (GLM) · new models on Hugging FaceOfficialAI score62

    Z.ai releases GLM-4.7-Flash, a 30B-A3B MoE model for lightweight deployment

    AIZ.ai has released GLM-4.7-Flash, a 30B-A3B MoE model that it positions as the strongest model in the 30B class. The model reports SWE-bench Verified 59.2 and τ²-Bench 79.5, and supports local deployment through vLLM and SGLang.

    Why it matters: The source lists benchmark scores against Qwen3-30B-A3B-Thinking-2507 and GPT-OSS-20B, letting readers compare the 30B-class MoE model directly with its named rivals.

Jan 14

Jan 14Wed
  1. Black Forest Labs · new models on Hugging FaceOfficialAI score62

    Black Forest Labs releases FLUX.2 [klein] 4B image model under Apache 2.0

    AIBlack Forest Labs released FLUX.2 [klein] 4B, a 4 billion parameter model that unifies text-to-image generation and image editing with multi-reference support. The source says it runs on consumer GPUs such as the RTX 3090 or 4070 with about 13GB VRAM, and its open weights are available under the Apache 2.0 license.

    Why it matters: The source specifies a 4 billion parameter model running on about 13GB VRAM under Apache 2.0, which helps readers judge whether local image generation fits their hardware.

Jan 1

Jan 1Thu
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score75

    Moonshot AI releases open-source multimodal agent model Kimi K2.5

    AIMoonshot AI released Kimi K2.5, an open-source native multimodal agentic model built by continual pretraining on about 15 trillion mixed visual and text tokens. The model card reports a 1T-parameter Mixture-of-Experts architecture with 32B activated parameters and a 256K context length, and it lists benchmark results against GPT-5.2, Claude 4.5 Opus, Gemini 3 Pro, DeepSeek V3.2, and Qwen3-VL-235B-A22B-Thinking. Weights and code are released under a Modified MIT License, with API access on the Moonshot platform.

    Why it matters: The model card gives a full benchmark table against GPT-5.2, Claude 4.5 Opus, and Gemini 3 Pro, useful for comparing open multimodal agent models.

Dec 20, 2025

Dec 20, 2025Sat
  1. MiniMax · new models on Hugging FaceOfficialAI score74

    MiniMax-M2.1 open-sources weights for coding and agent tasks

    AIMiniMax has released MiniMax-M2.1 model weights on Hugging Face, with API access on the MiniMax Open Platform and the MiniMax Agent product. The company reports gains over M2 on coding and agent benchmarks such as SWE-bench Verified (74.0) and VIBE average (88.6), and says it outperforms Claude Sonnet 4.5 on multilingual scenarios.

    Why it matters: The release pairs open weights with a broad benchmark table against Claude and GPT models, letting readers compare coding and agent claims directly.

Dec 16, 2025

Dec 16, 2025Tue
  1. Xiaomi MiMoOfficialAI score78

    Xiaomi releases open-source MiMo-V2-Flash MoE model for reasoning and coding

    AIXiaomi released and open-sourced MiMo-V2-Flash, a Mixture-of-Experts model with 309B total and 15B active parameters, under the MIT license. The company reports 73.4% on SWE-Bench Verified, the top score among open-source models, and inference at 150 tokens per second for $0.1 per million input tokens and $0.3 per million output tokens. It supports a hybrid thinking mode and a 256k context window.

    Why it matters: The post gives architecture, speculative decoding speedup, and pricing figures, which help readers judge how the efficiency claims are achieved and what they cost.

Dec 11, 2025

Dec 11, 2025Thu
  1. Nick TurleyXAI score78

    OpenAI introduces GPT-5.2 in ChatGPT for professional work

    AIOpenAI is introducing GPT-5.2 in ChatGPT, describing it as its most advanced model series for professional work. GPT-5.2 Thinking is positioned for tasks such as building spreadsheets and presentations, writing and reviewing production code, and analyzing long documents. The post says it beats or ties industry professionals on well-specified knowledge work tasks spanning 44 occupations 70.9% of the time on GDPval, and GPT-5.2 Instant, Thinking, and Pro begin rolling out to all tiers, starting with paid plans.

    Why it matters: The post links the model's professional-work focus to GDPval results across 44 occupations, showing how the claimed capability was measured.

    Image from @nickaturley's post
  2. Runway ResearchOfficialAI score62

    Runway Introduces GWM-1, a Real-Time General World Model Family

    AIRunway announced GWM-1, its first general world model family, built on Gen-4.5 and generating frames autoregressively in real time under interactive control. It comes in three variants: GWM Worlds for explorable environments, GWM Avatars for conversational characters, and GWM Robotics for robotic manipulation. Runway also says it is working toward unifying these domains under a single base world model, and GWM Robotics includes a Python SDK.

    Why it matters: The post separates three GWM-1 variants and ties each to a concrete use, which clarifies where a general world model would fit compared with a single model.

Nov 4, 2025

Nov 4, 2025Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score82

    Moonshot AI releases open-source Kimi K2 Thinking reasoning agent model

    AIMoonshot AI released Kimi K2 Thinking, an open-source thinking model that interleaves step-by-step reasoning with tool calls across 200 to 300 sequential invocations. The model is a 1T-parameter mixture-of-experts with 32B activated parameters and a 256k context window, and it uses native INT4 quantization for roughly 2x faster generation. The model card reports benchmark results on HLE, BrowseComp, and other tests, and recommends vLLM, SGLang, or KTransformers for deployment.

    Why it matters: The model card gives benchmark tables, quantization details, and deployment settings, letting readers compare Kimi K2 Thinking against GPT-5 and other models on specific tasks.

Nov 1, 2025

Nov 1, 2025Sat
  1. Runway ResearchOfficialAI score72

    Runway releases Gen-4.5, ranked first on the Text-to-Video benchmark

    AIRunway announced Gen-4.5, a video generation model that it says holds the top position on the Artificial Analysis Text-to-Video benchmark with 1,247 Elo points. The model is available across all paid Runway plans at comparable pricing, and the post lists limitations including causal reasoning errors, object permanence failures, and success bias.

    Why it matters: The post separates Runway's own ranking claim from the listed limitations, such as causal reasoning and object permanence errors, which helps judge where the model is reliable.

Oct 30, 2025

Oct 30, 2025Thu
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score60

    Moonshot AI releases Kimi Linear 48B hybrid linear attention models on Hugging Face

    AIMoonshot AI released Kimi Linear, a hybrid linear attention architecture with 48B total and 3B activated parameters and a 1M-token context length, on Hugging Face. The model card reports up to 6.3x faster TPOT than MLA at 1M tokens and up to 75% lower KV cache needs, and says it outperforms full attention on long-context and RL-style benchmarks.

    Why it matters: The model card gives concrete long-context speed and memory figures for a hybrid attention design, useful for judging whether linear attention can replace full attention in practice.

  2. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score72

    Moonshot AI releases Kimi Linear 48B-A3B hybrid attention models on Hugging Face

    AIMoonshot AI has released Kimi-Linear-Base and Kimi-Linear-Instruct, both 48B total and 3B activated parameters with a 1M context length, on Hugging Face. The models use Kimi Delta Attention in a 3:1 hybrid ratio with global MLA, cutting KV cache by up to 75% and boosting decoding throughput by up to 6x at 1M tokens. The KDA kernel is open-sourced in FLA, and the checkpoints were trained on 5.7T tokens.

    Why it matters: The model card gives concrete throughput and KV cache figures for a hybrid attention design, which helps readers weigh its long-context tradeoffs against full attention.

Oct 28, 2025

Oct 28, 2025Tue
  1. Cognition Blog (Devin, Windsurf)OfficialAI score72

    Cognition releases SWE-1.5, a coding agent model served at up to 950 tok/s

    AICognition has released SWE-1.5, a model optimized for software engineering that it says reaches near-frontier coding performance while running at up to 950 tok/s with Cerebras inference. The company reports it is 6x faster than Haiku 4.5 and 13x faster than Sonnet 4.5, and it is available now in Windsurf. The post's SWE-Bench Pro chart places SWE-1.5 at 40.08%, behind Sonnet 4.5 at 43.60%, and it notes that the model was trained with reinforcement learning on the Cascade agent harness.

    Why it matters: The post pairs a benchmark chart with a 950 tok/s speed claim and describes how harness, RL environments, and inference were co-designed, useful context for judging the speed-versus-quality tradeoff.