Skip to content

Formats · Latest news

Model releases

New models and updates: flagship releases, open weights, performance changes, and pricing changes.

183 top picks · 78 in the past 30 days · chosen from 1,047 items collected

Latest pick

Top picks archive · Page 3

Top picks 41–60 of 183

Sep 23

Sep 23Wed
  1. Google DeepMindAI score60

    Google DeepMind launches Gemini 3.8 Flash TTS and Flash-Lite TTS models

    AIGoogle DeepMind introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, text-to-speech models offering custom voice design, line-by-line performance control, and multilingual support across more than 100 languages. Flash TTS is rolling out to developers in the Gemini API and Google AI Studio and to everyone in Gemini Notebook, while Flash-Lite TTS is available to developers and in Google Vids. Voice replication requires consent verification, and generated audio carries SynthID watermarking.

    Why it matters: The source details the voice design, performance direction, and consent safeguards, showing how the model covers creative and high-volume use cases with access across several Google products.

  2. Baseten BlogAI score62

    Baseten launches NVIDIA Nemotron 3 Diarization with four latency profiles

    AIBaseten has made NVIDIA Nemotron 3 Diarization available as batch, streaming, and real-time diarized transcription presets. The single checkpoint serves four algorithmic latencies from 0.32 to 30.4 seconds, and the post reports DER of 9.8% on AISHELL-4 at the low profile versus 27.2% for Streaming Sortformer v2.1.

    Why it matters: The post shows one checkpoint serving four latency profiles with DER figures against named baselines, useful for judging real-time speaker labeling tradeoffs.

Sep 22

Sep 22Tue
  1. Fireworks AI BlogAI score65

    Fireworks releases Ember-1, a Kimi K3 variant that cuts reasoning tokens by about 40%

    AIFireworks Research released Ember-1, a specialized model built on Kimi K3 that it says delivers the same quality with 40% fewer tokens. Across five industry benchmarks, Ember-1 matched K3 max quality at a fraction of the cost, and in two customer A/B tests it used about 35% fewer tokens per task. It is available as a Research Preview on Serverless, and Fireworks is also launching training support for customized models.

    Why it matters: The source gives benchmark and A/B results for cutting reasoning tokens while holding quality, which bears on cost planning for coding and agent workloads.

  2. Tibor BlahoAI score88

    OpenAI launches GPT-6 Sol and Luna while Anthropic releases Claude Opus 5.5

    AIOpenAI released GPT-6 Sol and Luna, with API prices cut in half, while Anthropic released Claude Opus 5.5 at roughly Fable 5.1 level for 40% less than Opus 5. GPT-6 Sol and Luna cost $2/$10 and $0.10/$0.50 per million tokens, versus GPT-5.6 promotional pricing, and Opus 5.5 costs $4/$20 per million tokens. Sonnet 5.5 and Haiku 5.5 are announced for the coming weeks.

    Why it matters: The post links OpenAI's GPT-6 Sol and Luna pricing with Anthropic's Claude Opus 5.5 launch, which helps readers compare the two vendors' current frontier offerings.

  3. Greg BrockmanAI score81

    OpenAI launches GPT-6 Sol and Luna with 50% lower API prices than GPT-5.6

    AIOpenAI introduced GPT-6 Sol and GPT-6 Luna, which it says bring much of the strength of GPT-6 Astra into faster and more affordable models. The company also reports more efficient caching and inference, with API prices 50% lower than GPT-5.6 promotional pricing.

    Why it matters: The quoted announcement names specific pricing and access changes for Sol and Luna, which matter for teams weighing cost against the Astra tier.

  4. Noam BrownAI score78

    OpenAI releases GPT-6 Sol and Luna at 50% lower API prices

    AIOpenAI has released GPT-6 Sol and GPT-6 Luna, which it says build on GPT-6 Astra and offer faster, more affordable performance. API prices for Sol and Luna are 50% lower than GPT-5.6 promotional pricing, and Luna now costs $0.10 input and $0.50 output per 1M tokens. The author also notes an earlier 80% Luna price cut at the end of July, with output dropping from $6 to $0.50 within two months.

    Why it matters: The source gives concrete API price cuts across two model tiers, making the cost trend across recent releases easy to track for developers.

  5. Mike KriegerAI score67

    Anthropic launches Claude Opus 5.5, leading in coding and knowledge work

    AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. According to the quoted launch post, it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. The author says it leads in coding and knowledge work and praises its writing quality.

    Why it matters: The quoted launch post gives a concrete cost comparison, useful for weighing Opus 5.5 against earlier Opus and Fable 5.1 models for routine work.

  6. Sam BowmanAI score75

    Anthropic's Sam Bowman says Claude Opus 5.5 is safer, reducing misalignment risk

    AISam Bowman says Claude Opus 5.5 is sufficiently safer than its predecessors that releasing it more likely than not reduces misalignment risks. The quoted @claudeai post introduces Claude Opus 5.5 as the first model in the Claude 5.5 family, performing at the level of Claude Fable 5.1 on most tasks at 40% lower run cost than Opus 5.

    Why it matters: The post links a safety judgment to a model release, which is useful for readers weighing how Anthropic frames release decisions against misalignment risk.

  7. AnthropicAI score71

    Anthropic releases Claude Opus 5.5, the first model in its Claude 5.5 family

    AIAnthropic has made Claude Opus 5.5 available today, introducing it as the first model in its new Claude 5.5 family. According to the quoted @claudeai post, it performs at the level of Claude Fable 5.1 on most tasks and costs 40% less to run than Opus 5.

    Why it matters: The post gives a concrete cost comparison against Opus 5 and names the model family, helping readers gauge the trade-off between price and performance.

  8. Black Forest Labs · new models on Hugging FaceAI score62

    Black Forest Labs releases FLUX 3 Action, a 7B open-weights robot world action model

    AIBlack Forest Labs released FLUX 3 Action, an open-weights 7B world action model that outputs robot joint commands from camera frames, robot state, and a text instruction. On the RoboLab-120 benchmark it reports 42.92% task success, ahead of Cosmos3-Nano-Policy at 36.8% and π0.5 at 28.0%. The model is fine-tuned on DROID, is distributed under the FLUX Kommunity License v.1.0, and runs in about 32 GB of GPU memory in bfloat16.

    Why it matters: The model card gives a benchmark comparison, parameter counts, and an action contract, so readers can judge how it compares with existing robot policies.

  9. Black Forest Labs · new models on Hugging FaceAI score60

    Black Forest Labs releases FLUX 3 Action base weights for robot adaptation

    AIBlack Forest Labs has released flux-3-action-base, an open-weights 7B world action model that takes camera frames, robot state, and a text instruction to output the next action chunk. The release is an adaptation component rather than a complete robot policy, and new embodiments require their own action heads. The source says the weights are paired with shared video VAE and Qwen3-VL-4B-Instruct text encoders and is governed by the FLUX Kommunity License v.1.0.

    Why it matters: The source separates the adaptation base from full robot policies and states the shared encoders and new-embodiment requirements, which clarifies what developers must still build for their robots.

  10. METR BlogAI score62

    METR's preliminary evaluation finds Claude Opus 5.5 is an incremental AI R&D gain over Fable 5.1

    AIMETR's preliminary evaluation concludes that Claude Opus 5.5 likely gives slightly higher AI R&D productivity uplift than Fable 5.1 but is unlikely to fully automate AI R&D. The evaluation used five capability tasks over 10 business days of API access, and METR says Anthropic reviewed and edited the summary before sign-off.

    Why it matters: The report separates two claims about AI R&D acceleration and discloses that Anthropic reviewed the summary, which helps readers weigh its independence and evidence.

  11. Gemini API ChangelogAI score62

    Gemini 3.8 Flash TTS and Flash-Lite TTS become generally available with a new Voices endpoint

    AIGoogle made the Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS models generally available, along with the Gemini API Voices endpoint. Flash TTS is positioned for studio-grade voice fidelity and long-form multi-turn stability, while Flash-Lite TTS targets high-throughput, real-time voice agents and replaces gemini-3.1-flash-tts-preview. The update adds voice design, voice replication with consent verification, and access to 150+ prebuilt and custom voices.

Sep 21

Sep 21Mon
  1. Tencent HunyuanAI score67

    Tencent Hy4 preview compressed to 214 GiB with mixed-precision quantization

    AITencent Hunyuan says it shrank the 770B-parameter Hy4 preview from roughly 1.5TB to 214 GiB while keeping the parameter count unchanged. The quoted Zhihu post by a Tencent Hunyuan quantization team member describes the method: a 1.25-bit sparse ternary encoding, mixed precision across expert layers, and STQ1_0 CUDA kernels in llama.cpp. The author reports nearly unchanged MRCR retrieval and a small decline in math.

    Why it matters: The quoted Zhihu post explains how Hy4 preview's weights were quantized and kept usable at inference, a concrete engineering case for compressing large MoE models.

  2. Claude Apps Release NotesAI score62

    Anthropic launches Claude Opus 5.5, first model in its 5.5 family

    AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. The company says it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

    Why it matters: The release note gives a direct comparison to Claude Fable 5.1 and a 40% running cost reduction versus Opus 5, which helps readers weigh the tradeoff.

  3. Xiaomi MiMoAI score67

    Xiaomi MiMo open-sources Pro, Flash, and a 9B distilled model

    AIXiaomi MiMo announced open-source releases of Pro and Flash, the MiMo-V2.6-Distill-Qwen-9B model, a technical report, over 7K RL task environments, an end-to-end RL framework, and composable mini-harnesses. The attached table shows MiMo-V2.6-Distill-Qwen-9B after SFT and after RL compared with Qwen3.5-9B, with RL scores higher on most listed benchmarks, such as SWE-bench Verified at 66.2 versus 60.0.

    Why it matters: The table compares a 9B distilled model against Qwen3.5-9B on coding, cyber, and agent benchmarks, showing how the reinforcement learning stage changes results.

  4. Xiaomi MiMoAI score78

    Xiaomi releases open-weight MiMo-V2.6 Pro and Flash omnimodal models

    AIXiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.

    Why it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

  5. Xiaomi MiMo · new models on Hugging FaceAI score67

    Xiaomi releases MiMo-V2.6-Flash-RL, a 309B sparse MoE model with 1M context

    AIXiaomi released MiMo-V2.6-Flash-RL, an efficiency-balanced checkpoint in its MiMo-V2.6 series, on Hugging Face. The model is a sparse MoE with 309B total and 15B activated parameters, supports text, image, video, and audio input, and offers a 1M-token context. The technical report says it was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs its benchmark tables with the RL training method, which helps readers judge how the checkpoint's scores relate to its training approach.

  6. Xiaomi MiMo · new models on Hugging FaceAI score74

    Xiaomi MiMo-V2.6-Pro-RL released as 1.02T-parameter omnimodal model

    AIXiaomi MiMo released MiMo-V2.6-Pro-RL on Hugging Face, a sparse MoE model with 1.02T total and 42B activated parameters and a 1M-token context. The technical report says it accepts text, image, video, and audio, and was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs a 1.02T-parameter MoE model with an RL-based self-improvement method, useful for judging how reinforcement learning is scaled in frontier open models.