Skip to content

Formats · Latest news

Model releases

New models and updates: flagship releases, open weights, performance changes, and pricing changes.

183 top picks · 78 in the past 30 days · chosen from 1,048 items collected

Latest pick

Top picks archive · Page 5

Top picks 81–100 of 183

Sep 6

Sep 6Sun
  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score62

    OpenBMB releases MiniCPM5-2B, a 2B open-source model with open training data

    AIOpenBMB has released MiniCPM5-2B, a dense 2B Transformer built for on-device and resource-constrained deployment, with an average score of 53.9 in its comparison set. The release also opens the UltraData datasets behind it, including UltraX, UltraData-Code, UltraData-SFT-Agent-2609 and UltraData-RL-2609, and includes GGUF, MLX, GPTQ and DSpark variants for common runtimes.

    Why it matters: The release pairs a 2B model with open training datasets and reports per-benchmark comparisons against named same-size and larger models, letting readers check the claims directly.

Sep 3

Sep 3Thu
  1. Mark ChenAI score80

    Mark Chen announces GPT-6 Astra with computer use and agent oversight

    AIOpenAI researcher Mark Chen announced GPT-6 Astra, which he described as the company's most capable and aligned model yet. He said it can build and test software, work across apps on a computer, and help with open scientific problems. The post also highlights improved computer use compared with Operator and stronger monitoring that can stop potentially unauthorized agent actions.

    Why it matters: The post links a named model release to specific capabilities like computer use and aligned agent behavior, giving readers concrete claims to check against the model.

  2. Google DeepMind · The KeywordAI score72

    Google DeepMind releases WeatherNext 3, a global weather model with hourly satellite-based forecasts

    AIGoogle DeepMind and Google Research introduced WeatherNext 3, which generates hourly global forecasts at up to 5-kilometer resolution using live geostationary satellite data. The company reports that precipitation forecasts improved by up to 60% against IMERG in medium-range evaluations, and that longer-range precipitation forecasts are up to 50% more accurate. The model is now available across Search, Gemini, Google Maps, Google Maps Platform Weather API, Google Earth Engine, BigQuery, and Google Cloud Storage.

    Why it matters: The post explains how training on live satellite data and station observations changes resolution and update frequency, with precipitation accuracy gains reported against named baselines.

  3. Google DeepMind · YouTubeAI score72

    Google DeepMind's WeatherNext 3 offers hourly, 5km-resolution weather forecasts

    AIGoogle DeepMind introduced WeatherNext 3, a weather forecasting model that learns directly from satellite feeds and ground-level weather station data. It produces a fresh forecast every hour, compared with the six-hour refresh typical of traditional models, with native 5km resolution for temperature and humidity. It is available through Google Search, Gemini, Google Maps and more.

    Why it matters: The source shows a shift from six-hourly to hourly refresh and 5km local resolution, which matters for energy planning and local forecasting.

Sep 2

Sep 2Wed
  1. NVIDIA · new models on Hugging FaceAI score67

    NVIDIA releases Nemotron-3-Labs-Ultra-Math-RL for mathematical proof reasoning

    AINVIDIA has published Nemotron-3-Labs-Ultra-Math-RL on Hugging Face, a 550B total, 55B active parameter model for solving difficult math problems and identifying proof mistakes. The model is part of an ensemble that reached gold-medal level at the International Mathematical Olympiad 2026, and it is available for commercial and non-commercial use under the OpenMDW-1.1 license. Deployment is designed for NVIDIA Blackwell or Hopper GPUs, with a recommended minimum of 8× B200 on a single node and a context length of up to 1M tokens.

    Why it matters: The release details the model's math-proof role, its 550B total and 55B active parameters, and its vLLM deployment requirements for teams weighing adoption.

  2. Google AI StudioAI score78

    Google releases Gemini 3.8 Flash and restricted 3.8 Flash Cyber model

    AIGoogle introduces Gemini 3.8 Flash for coding, agentic tasks, and multi-step reasoning, priced at $0.75 per million input tokens and $3.75 per million output tokens during the introductory period. Gemini 3.8 Flash Cyber targets vulnerability detection and automated patching and is available only to trusted defenders through the new Fairwind Program. The introductory price expires December 31, 2026, after which $1.50 and $7.50 per million tokens apply.

    Why it matters: The post separates a general coding and agent model from a restricted cyber variant, showing how one shared core is deployed under different access and safety tiers.

Sep 1

Sep 1Tue
  1. Anthropic · YouTubeAI score78

    Anthropic releases Claude Fable 5.1, an upgrade to its most capable model class

    AIAnthropic has released Claude Fable 5.1, the latest upgrade to its most capable class of models, and it is available everywhere today. The company says it handles complex, long-running, multi-step work and avoids shortcuts when fixing root causes of software issues. At lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost, according to Anthropic's benchmarks.

    Why it matters: The source names the upgraded model class and its cost tradeoff at lower effort levels, which helps readers weigh it against the earlier version for their own workloads.

  2. Anthropic · YouTubeAI score72

    Anthropic releases Claude Fable 5.1 for complex, long-running tasks

    AIAnthropic has released Claude Fable 5.1, an upgrade to its most capable model class, and says it is available everywhere today. The company reports that at lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost. It is described as strong at complex multi-step work, such as long proofs and contracts with hundreds of cross-references, and at fixing root causes in software issues.

    Why it matters: The source reports cost and effort-level tradeoffs for long-running tasks, helping readers judge whether the upgrade changes their workloads or budgets.

  3. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score60

    Shanghai AI Lab releases Intern Lumina U2 unified multimodal model on Hugging Face

    AIShanghai AI Lab's InternLM has published Intern Lumina U2, a 16B-parameter MoE model with 1B active parameters that handles text QA, image generation and editing, and image, video, and 3D understanding. The model uses an 8-codebook fully-discrete visual representation built on AToken. Checkpoints are provided for Huawei Ascend NPUs and NVIDIA GPUs under Apache 2.0, with the technical report still listed as coming soon.

    Why it matters: The model unifies text, image, video, and 3D understanding with image generation in one framework, a broader scope than single-modality releases.

Aug 31

Aug 31Mon
  1. Claude Apps Release NotesAI score72

    Anthropic launches Claude Fable 5.1 and Claude Mythos 5.1 models

    AIAnthropic has launched Claude Fable 5.1 and Claude Mythos 5.1, which it describes as the world's most advanced models for coding and knowledge work. The release notes link to a blog post with more details, but the notes themselves give no benchmarks or specifications.

    Why it matters: The source names two new model versions and points to a companion blog post, so readers can compare the release details there.

  2. DeepSeek · new models on Hugging FaceAI score65

    DeepSeek releases V4-Flash-Vision-Exp, an experimental multimodal agent model

    AIDeepSeek introduces DeepSeek-V4-Flash-Vision-Exp, its first experimental multimodal model in the DeepSeek-V4 family, built on V4-Flash with visual modules. It reports substantial gains over DeepSeek-V4-Flash-0731 on multimodal agent benchmarks, such as ApexBench at 36.5 versus 26.2, while keeping text agent performance comparable. The repository provides tokenizer files, prompt encoding, vLLM and SGLang serving instructions, and is licensed under MIT.

    Why it matters: The source compares the model with its text-only predecessor and Opus-4.8 on agent benchmarks, showing where vision gains occur and where text performance holds.

Aug 28

Aug 28Fri
  1. Unsloth AIAI score70

    Unsloth shows how to run GLM-5.3 locally with 2-bit quantization

    AIUnsloth AI published a guide for running GLM-5.3 locally using quantized GGUF weights. The 2-bit version is reduced from 1.51TB to 239GB and retains about 81% accuracy, and it can run on a 256GB Mac or RAM/VRAM setups.

    Why it matters: The guide shows which quantization levels fit local memory budgets and how much accuracy each costs, useful for planning a local deployment.

Aug 27

Aug 27Thu
  1. Unsloth AIAI score70

    GLM-5.3-Flash can run locally with Unsloth GGUF quantization on 128GB RAM

    AIUnsloth says GLM-5.3-Flash can run locally, with a 3-bit GGUF version needing 128GB of RAM and the 1-bit version working on 102GB of RAM or VRAM. The guide's table lists memory needs from 100GB at 1-bit to 650GB at BF16, and reports that the 1-bit quant keeps 71% of top-1% accuracy while being 85% smaller than BF16.

    Why it matters: The guide gives concrete memory requirements for each quantization level, which helps readers judge whether the model fits their hardware.

  2. OpenBMB (MiniCPM) · new models on Hugging FaceAI score65

    OpenBMB releases MiniCPM5-2B-SFT, a 2B open model with SFT-only checkpoint

    AIOpenBMB released MiniCPM5-2B-SFT, an SFT-only BF16 checkpoint taken before RL and OPD, within its MiniCPM5-2B series. The model is a 2B dense Transformer built for on-device and local deployment, with 131,072-token context and the same training recipe as the final release.

    Why it matters: The source gives concrete benchmark averages against same-size and larger models, plus released training data and multiple deployment formats, useful for judging a compact on-device model.

  3. Tencent · new models on Hugging FaceAI score80

    Tencent open-sources Hy4 preview, a 770B-parameter MoE model

    AITencent's Hy Team released Hy4 preview, a Mixture-of-Experts model with 770B total parameters and 49B activated per token, with a 1M context length. Hugging Face hosts the Instruct model and an FP8 quantized version under the Apache License 2.0, with vLLM and SGLang deployment instructions provided.

    Why it matters: The model card gives architecture, activated parameters, and vLLM and SGLang deployment recipes, useful for judging whether the release fits your serving setup.

  4. Qwen · new models on Hugging FaceAI score62

    Qwen-Drive-1.0 releases open weights for driving VQA, perception, and planning

    AIQwen has published Qwen-Drive-1.0-4B on Hugging Face, a vision-language model for autonomous driving built on Qwen3.5-4B. The release includes a BEV perception head and two Planning Experts, planner-sft and planner-rl, with code and an inference example in the linked GitHub repository.

    Why it matters: The source gives concrete benchmark results and a runnable setup, letting readers judge how a driving VLM with planning and perception heads compares with existing systems.

Aug 26

Aug 26Wed
  1. LMSYS OrgAI score65

    Zhipu's GLM-5.3-Flash adds native vision with day-0 SGLang support

    AIZ.ai released GLM-5.3-Flash, a 320B-A18B model, with day-0 support in SGLang, after appearing earlier as ox-alpha. The post calls it the first native multimodal model in the GLM-5 series and says it outperforms GLM-5.2 at one-tenth the cost, with stable 1M-token long-context performance.

    Why it matters: The post reports GLM-5.3-Flash's native multimodal design, its efficiency claims, and day-0 SGLang support, which bear on running it in practice.

  2. Unsloth AIAI score78

    Unsloth explains how to run Qwen3.8-Flash-Next locally on 75GB RAM

    AIUnsloth announces that Qwen3.8-Flash-Next can be run locally through its GGUF quantizations. The source says the 1-bit version needs 75GB of RAM or unified memory, and that the 125B MoE model is reported to outperform Claude-Opus-4.6 (Max).

    Why it matters: The source gives concrete local hardware requirements, quantization sizes, and a guide, showing how a 125B MoE model can run on a 75GB RAM setup.

Aug 25

Aug 25Tue
  1. Z.ai Release NotesAI score62

    Z.ai releases GLM-5.3-Flash with native visual capabilities and hybrid architecture

    AIZ.ai has released GLM-5.3-Flash, a model with native visual capabilities that observe interfaces, rendering results, and interaction feedback across code, browsers, and GUIs. It uses a hybrid linear and sparse attention architecture with 320B total parameters and 18B activated, which the company says significantly reduces compute and KV-cache requirements. The release notes also describe support for office document and financial research workflows.

    Why it matters: The release notes give GLM-5.3-Flash's architecture, parameter counts, and cybersecurity findings, which make the model's scope concrete for comparison with earlier GLM releases.