Skip to contentSkip to stories

Updated

#Model release

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 23

Sep 23Wed
  1. Google Developers BlogOfficialAI score62

    Google reproduces Olmo 3 7B pre-training in MaxText on TPUs

    AIGoogle Developers reproduced Ai2's Olmo 3 7B from scratch in MaxText on Google Cloud TPUs, covering both the stage-1 pre-training run and the stage-2 mid-training anneal. The match was checked on held-out C4 loss, an 8-task accuracy suite, multi-domain perplexity, and token-level KL, not just the training loss curve. The post also describes a data-loader bug that made training loss look better than the reference while held-out metrics did not move.

    Why it matters: The post documents how a faithful reproduction was verified on held-out metrics, including a data bug that training loss alone would have hidden.

  2. Microsoft CopilotOfficialAI score62

    Claude Opus 5.5 and GPT-6 Sol roll out in Microsoft Copilot apps

    AIAnthropic's Claude Opus 5.5 and OpenAI's GPT-6 Sol are rolling out in Microsoft Copilot across Word, Excel, PowerPoint, Chat, Cowork, and Copilot Studio. The post says Work IQ supplies work context, so users can choose the model that fits each task.

    Why it matters: The post shows where two newly named models become selectable inside Microsoft Copilot's Office apps and Studio, which matters for users choosing a model per task.

  3. Black Forest LabsOfficialAI score40

    Black Forest Labs unveils FLUX 3 Action for robotic control

    AIBlack Forest Labs says FLUX 3 Action takes recent camera frames, the system's current state, and a task description to return the next 32 actions. The model also predicts how the scene will change, and runs a continuous observe-plan-act-adjust loop. It can recover from its own mistakes, according to the company.

    Video from @bfl_ai's post
  4. Black Forest LabsOfficialAI score40

    FLUX 3 Action uses a smaller architecture to predict actions and frames

    AIBlack Forest Labs says FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3 but uses a smaller architecture. The company attributes this smaller size to more efficient representations learned through its Self-Flow research. During midtraining, the model was trained to predict actions and future frames together.

    Video from @bfl_ai's post
  5. Black Forest LabsOfficialAI score62

    Black Forest Labs releases FLUX 3 Action, a robot policy model

    AIBlack Forest Labs says its FLUX 3 Action, a single-step 7B checkpoint, outperforms every other open policy on RoboLab. It processes each second of robot motion 1.45× to 1.66× faster than Pi0.5, and uses a 2.13-second action horizon versus Pi0.5's 1 second. The company adds that its guidance-distilled checkpoint raises the state-of-the-art RoboLab success rate while running 2.85× to 3.15× faster than the previous leading open WAM.

    Why it matters: The post pairs a success-rate gain on RoboLab with measured speed against Pi0.5, showing whether a robot policy can match video-prediction accuracy without losing responsiveness.

    Image from @bfl_ai's post
  6. Black Forest LabsOfficialAI score67

    Black Forest Labs releases FLUX 3 Action, an open 7B world action model for robots

    AIBlack Forest Labs says FLUX 3 Action is an open-weights 7B world action model that ranks first on the RoboLab benchmark. The company says it outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster. The model predicts video and actions together, and the company is releasing the weights, code, fine-tuning recipe, benchmarks, and examples. It also integrated the model into Hugging Face's LeRobot with NVIDIA, with edge deployment on NVIDIA Jetson.

    Why it matters: The release pairs benchmark results with the trade-off it claims to remove between world action model performance and VLA speed, which is useful context for robotics teams weighing open models.

    Video from @bfl_ai's post
  7. Google AI StudioOfficialAI score62

    Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle introduces Gemini 3.8 Flash TTS for creative voice design and Gemini 3.8 Flash-Lite TTS for high-volume, cost-efficient speech generation. Flash TTS supports voice creation from natural language prompts across more than 100 languages and dialects, and both models are rolling out today in the Gemini API and Google AI Studio, with enterprise access coming soon via Gemini Enterprise.

    Why it matters: The source details two new TTS models, their voice design and replication features, and their rollout across developer, enterprise, and consumer channels, useful for comparing voice model options.

  8. Logan KilpatrickXAI score33

    Google releases Gemini 3.8 Flash TTS with lower cost than 3.1

    AIGoogle has released Gemini 3.8 Flash TTS, a text-to-speech model that costs less than its previous 3.1 Flash TTS model. The new model is accompanied by a new experience in Google AI Studio.

  9. Logan KilpatrickXAI score46

    Google launches Gemini 3.8 Flash and Flash-Lite TTS text-to-speech models

    AIGoogle introduced Gemini 3.8 Flash and Flash-Lite TTS, new state-of-the-art text-to-speech models. They add a new voice design experience, over 2,000 production-ready voices, voice replication, and support for 100 languages. The models rank first on Hume AI's voice benchmarks, and voice remixing is coming soon.

  10. Google for DevelopersOfficialAI score52

    Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle announced two new Gemini 3.8 text-to-speech models, positioned as its most expressive yet. Gemini 3.8 Flash TTS targets creative work, letting developers use natural language to define vocal personas, cues, pacing, and dialects, while Gemini 3.8 Flash-Lite TTS is built for high-volume pipelines such as bulk audiobook production and audio dubbing. Both are available now through the Gemini API in Google AI Studio.

  11. Google AIOfficialAI score40

    Google rolls out Gemini 3.8 Flash TTS and Flash-Lite TTS models

    AIGoogle is rolling out two Gemini 3.8 text-to-speech models starting today: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Developers can access both in Google AI Studio and the Gemini API, while consumers get Flash TTS in NotebookLM and Flash-Lite TTS in Google Vids. Both models are coming soon to Gemini Enterprise.

  12. Google AIOfficialAI score62

    Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle AI launched Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, which it describes as its most expressive audio models yet. The models support custom voices across 100+ languages, 2,000+ prebuilt voices, multi-speaker conversations, line-by-line delivery control, and cues such as <laughs> and |mhm|. Google positions Flash TTS for bespoke voices in gaming, audiobooks, and podcasts, and Flash-Lite TTS for near real-time voice agents, high-volume dubbing, and bulk audio.

    Why it matters: The post separates a high-fidelity creative model from a cost-efficient model for real-time agents and bulk dubbing, clarifying which workload each one targets.

    Video from @GoogleAI's post
  13. Google DeepMindOfficialAI score44

    Google launches Gemini 3.8 Flash and Flash-Lite TTS models for custom audio

    AIGoogle DeepMind has introduced two text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Gemini 3.8 Flash TTS lets users design unique voices with distinct accents and characteristics, while Flash-Lite TTS is built for efficiency and scale, offering user-created styles or a production-ready voice library.

    Video from @GoogleDeepMind's post
  14. Google AI StudioOfficialAI score42

    Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS audio models

    AIGoogle introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, describing them as its most expressive audio generation models yet. The models are designed to help creators, developers, and enterprises build richer, more expressive audio experiences, and can be tried through the Gemini API and AI Studio.

    Video from @GoogleAIStudio's post
  15. Google DeepMind · YouTubeOfficialAI score46

    Gemini 3.8 text-to-speech lets developers design and clone custom voices

    AIGoogle DeepMind's latest Gemini Audio models let developers design new vocal personas from natural language prompts, directing pacing, back channeling, and dialect shifts line by line. Developers can also recreate consistent adult voice profiles from a 30-second audio sample, with built-in consent verification, SynthID watermarking, and C2PA credentials.

  16. Baseten BlogOfficialAI score62

    Baseten launches NVIDIA Nemotron 3 Diarization with four latency profiles

    AIBaseten has made NVIDIA Nemotron 3 Diarization available as batch, streaming, and real-time diarized transcription presets. The single checkpoint serves four algorithmic latencies from 0.32 to 30.4 seconds, and the post reports DER of 9.8% on AISHELL-4 at the low profile versus 27.2% for Streaming Sortformer v2.1.

    Why it matters: The post shows one checkpoint serving four latency profiles with DER figures against named baselines, useful for judging real-time speaker labeling tradeoffs.

  17. ModelScopeOfficialAI score44

    NVIDIA releases Nemotron 3 Diarization for live speaker attribution

    AINVIDIA's Nemotron 3 Diarization is now available on ModelScope, labeling speakers and timestamps in streaming audio for up to eight speaker slots per conversation. The 99.2M-parameter model uses an end-to-end streaming architecture built on NVIDIA's Streaming Sortformer, running on Ampere, Hopper, and Blackwell GPUs via NeMo Speech C++. It is designed to pair with existing ASR systems such as Nemotron ASR, Parakeet, Canary, or Whisper to produce speaker-attributed transcripts.

    Image from @ModelScope2022's post
  18. ModelScopeOfficialAI score40

    TeleOCR: 1.2B vision-language model parses documents, tops OmniDocBench v1.6

    AITeleOCR, a lightweight 1.2B vision-language model released under Apache 2.0, parses digital PDFs and warped phone photos without a separate dewarping model. It scores 96.87 overall on OmniDocBench v1.6, the highest among listed specialized VLMs, and ranks #1 in the ICDAR 2026 Sci-ImageMiner Challenge. It supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.

    Image from @ModelScope2022's post
  19. howie.seriousXAI score22

    Opus 5.5 praised for language, visual taste, code quality, and token efficiency

    AIThe X user howie.serious says Claude Opus 5.5 delivers high language quality, good visual taste, strong code quality, and notably low token usage. He also compares Anthropic's reported roughly $2 trillion IPO valuation with OpenAI's roughly $1.2 trillion fundraising valuation, arguing OpenAI is worth about 0.6 Anthropics and the gap may widen.

  20. QwenOfficialAI score62

    Qwen-Audio-3.1 upgrades ASR, TTS and Realtime and adds two new models

    AIAlibaba's Qwen team released Qwen-Audio-3.1, upgrading its ASR, TTS and Realtime models and adding TTS-Next and ASR-Next. The post says TTS prices fell about 70%, Realtime about 85%, and ASR up to 95%. More APIs are coming soon.

    Why it matters: The post lists the new audio models and price changes together, which helps developers compare the upgraded lineup with their current speech workflows.

    Image from @Alibaba_Qwen's post
  21. ModelScopeOfficialAI score62

    Xiaomi MiMo-V2.6 open-sourced as a multimodal agent model family under MIT License

    AIXiaomi has released MiMo-V2.6 as an open model family under the MIT License, designed for large-scale reinforcement learning. MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index, with 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified. The 1.02T-parameter MoE activates 42B parameters and supports text, image, video, and audio input with a 1M-token context.

    Why it matters: The post links benchmark results, parameter scale, and a multi-agent RL training run, giving readers concrete figures to compare against other open models.

    Image from @ModelScope2022's post
  22. KrASIA · Big TechNewsAI score46

    Tencent Hy Image 3.5 preview refined through its consumer and business products

    AITencent has released a preview of its Hy Image 3.5 image generation model, which product teams across Yuanbao, WorkRally, Ima, and other services are helping refine through co-design. Tencent Cloud prices the model at USD 0.024 per 2K output image, and it supports text-to-image and image-to-image generation with up to five reference images. Tencent said an internal blind evaluation found it on par with ByteDance's Seedream 5.0 Pro and slightly better than Nano-Banana Pro and Qwen-Image-3.0 Pro.

  23. swyxXAI score29

    Opus 5.5 becomes new default for AINews, writing more concise

    AIswyx reports that Claude Opus 5.5 is now the default model for Latent Space's AINews, after side-by-side testing against Sol showed more concise and tasteful reporting with less "slopese" than Opus 5. The post links to the AINews issue titled "Claude Opus 5.5: The New Default."

    Image from @swyx's post
  24. ModelScopeOfficialAI score62

    Shanghai AI Lab and SJTU release open-weight 8.9B NCP-ArchPreview model under Apache 2.0

    AIShanghai AI Lab and SJTU's LUMIA Lab released NCP-ArchPreview, an 8.9B open-weight language model under Apache 2.0. The model reportedly reaches OLMo-3-7B's final Stage 1 loss using 51.3% of the tokens from the 5.73T Dolma 3 corpus, a 1.95× convergence gain. Its concept module jointly predicts tokens and concepts, and domain adaptation updates only its 17M parameters while the token backbone stays frozen.

    Why it matters: The post pairs an Apache 2.0 open-weight release with training-efficiency figures, showing how the concept module adapts to new domains with few trainable parameters.

    Image from @ModelScope2022's post

Sep 22

Sep 22Tue
  1. ModelScopeOfficialAI score62

    inclusionAI open-sources Ming-Image-0.1-Design models for visual design

    AIinclusionAI open-sources the Ming-Image-0.1-Design family, two complementary 6B models for visual-design workflows, under an MIT License. Design generates complete UIs, dashboards, infographics, and posters up to 2048×2048 with native transparent RGBA output, and Layer decomposes flattened graphics into independently editable RGBA layers.

    Why it matters: The post separates a design-generation model from a layer-decomposition model, letting users compare two distinct visual-design workflows under one MIT license.

    Image from @ModelScope2022's post
  2. Fireworks AI BlogOfficialAI score65

    Fireworks releases Ember-1, a Kimi K3 variant that cuts reasoning tokens by about 40%

    AIFireworks Research released Ember-1, a specialized model built on Kimi K3 that it says delivers the same quality with 40% fewer tokens. Across five industry benchmarks, Ember-1 matched K3 max quality at a fraction of the cost, and in two customer A/B tests it used about 35% fewer tokens per task. It is available as a Research Preview on Serverless, and Fireworks is also launching training support for customized models.

    Why it matters: The source gives benchmark and A/B results for cutting reasoning tokens while holding quality, which bears on cost planning for coding and agent workloads.

  3. Tibor BlahoXAI score88

    OpenAI launches GPT-6 Sol and Luna while Anthropic releases Claude Opus 5.5

    AIOpenAI released GPT-6 Sol and Luna, with API prices cut in half, while Anthropic released Claude Opus 5.5 at roughly Fable 5.1 level for 40% less than Opus 5. GPT-6 Sol and Luna cost $2/$10 and $0.10/$0.50 per million tokens, versus GPT-5.6 promotional pricing, and Opus 5.5 costs $4/$20 per million tokens. Sonnet 5.5 and Haiku 5.5 are announced for the coming weeks.

    Why it matters: The post links OpenAI's GPT-6 Sol and Luna pricing with Anthropic's Claude Opus 5.5 launch, which helps readers compare the two vendors' current frontier offerings.

    Image from @btibor91's post
  4. Tri DaoXAI score44

    Rigel: 2.3B hybrid Mamba-2 MoE nears Llama-3.2-3B with <1% FLOPs

    AIMayank's Rigel, a 2.3B-parameter MoE (360M active) hybrid Mamba-2 model, was pretrained across H100, A100, V100 GPUs and TPU v5p/v6e on one codebase. The model lands within a few points of Llama-3.2-3B while using under 1% of its pretraining FLOPs. Tri Dao praised the work's engineering effort and the model's strength for its small size.

  5. Simon WillisonXAI score60

    OpenAI launches GPT-6 Sol and Luna with 50% lower API prices

    AIOpenAI introduced GPT-6 Sol and GPT-6 Luna, which build on advances behind GPT-6 Astra. The company says it cut API prices 50% for Sol and Luna compared with GPT-5.6 promotional pricing, passing on caching and inference efficiency gains. Simon Willison notes GPT-6 Luna costs half of GPT-5.6 Luna and calls Luna his favorite model for building product features because of its cost and speed.

  6. Greg BrockmanXAI score81

    OpenAI launches GPT-6 Sol and Luna with 50% lower API prices than GPT-5.6

    AIOpenAI introduced GPT-6 Sol and GPT-6 Luna, which it says bring much of the strength of GPT-6 Astra into faster and more affordable models. The company also reports more efficient caching and inference, with API prices 50% lower than GPT-5.6 promotional pricing.

    Why it matters: The quoted announcement names specific pricing and access changes for Sol and Luna, which matter for teams weighing cost against the Astra tier.