Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 23

Sep 23Wed
  1. Google AI StudioAI score62

    Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle introduces Gemini 3.8 Flash TTS for creative voice design and Gemini 3.8 Flash-Lite TTS for high-volume, cost-efficient speech generation. Flash TTS supports voice creation from natural language prompts across more than 100 languages and dialects, and both models are rolling out today in the Gemini API and Google AI Studio, with enterprise access coming soon via Gemini Enterprise.

  2. Google for DevelopersAI score52

    Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle announced two new Gemini 3.8 text-to-speech models, positioned as its most expressive yet. Gemini 3.8 Flash TTS targets creative work, letting developers use natural language to define vocal personas, cues, pacing, and dialects, while Gemini 3.8 Flash-Lite TTS is built for high-volume pipelines such as bulk audiobook production and audio dubbing. Both are available now through the Gemini API in Google AI Studio.

  3. Google AIAI score62

    Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle AI launched Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, which it describes as its most expressive audio models yet. The models support custom voices across 100+ languages, 2,000+ prebuilt voices, multi-speaker conversations, line-by-line delivery control, and cues such as <laughs> and |mhm|. Google positions Flash TTS for bespoke voices in gaming, audiobooks, and podcasts, and Flash-Lite TTS for near real-time voice agents, high-volume dubbing, and bulk audio.

    Video from @GoogleAI's post
  4. Google DeepMindAI score60

    Google DeepMind launches Gemini 3.8 Flash TTS and Flash-Lite TTS models

    AIGoogle DeepMind introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, text-to-speech models offering custom voice design, line-by-line performance control, and multilingual support across more than 100 languages. Flash TTS is rolling out to developers in the Gemini API and Google AI Studio and to everyone in Gemini Notebook, while Flash-Lite TTS is available to developers and in Google Vids. Voice replication requires consent verification, and generated audio carries SynthID watermarking.

    Why it matters: The source details the voice design, performance direction, and consent safeguards, showing how the model covers creative and high-volume use cases with access across several Google products.

  5. Google DeepMind · YouTubeAI score46

    Gemini 3.8 text-to-speech lets developers design and clone custom voices

    AIGoogle DeepMind's latest Gemini Audio models let developers design new vocal personas from natural language prompts, directing pacing, back channeling, and dialect shifts line by line. Developers can also recreate consistent adult voice profiles from a 30-second audio sample, with built-in consent verification, SynthID watermarking, and C2PA credentials.

  6. ModelScopeAI score44

    NVIDIA releases Nemotron 3 Diarization for live speaker attribution

    AINVIDIA's Nemotron 3 Diarization is now available on ModelScope, labeling speakers and timestamps in streaming audio for up to eight speaker slots per conversation. The 99.2M-parameter model uses an end-to-end streaming architecture built on NVIDIA's Streaming Sortformer, running on Ampere, Hopper, and Blackwell GPUs via NeMo Speech C++. It is designed to pair with existing ASR systems such as Nemotron ASR, Parakeet, Canary, or Whisper to produce speaker-attributed transcripts.

    Image from @ModelScope2022's post
  7. ModelScopeAI score40

    TeleOCR: 1.2B vision-language model parses documents, tops OmniDocBench v1.6

    AITeleOCR, a lightweight 1.2B vision-language model released under Apache 2.0, parses digital PDFs and warped phone photos without a separate dewarping model. It scores 96.87 overall on OmniDocBench v1.6, the highest among listed specialized VLMs, and ranks #1 in the ICDAR 2026 Sci-ImageMiner Challenge. It supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.

    Image from @ModelScope2022's post
  8. ModelScopeAI score62

    Xiaomi MiMo-V2.6 open-sourced as a multimodal agent model family under MIT License

    AIXiaomi has released MiMo-V2.6 as an open model family under the MIT License, designed for large-scale reinforcement learning. MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index, with 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified. The 1.02T-parameter MoE activates 42B parameters and supports text, image, video, and audio input with a 1M-token context.

    Image from @ModelScope2022's post
  9. KrASIA · Big TechAI score46

    Tencent Hy Image 3.5 preview refined through its consumer and business products

    AITencent has released a preview of its Hy Image 3.5 image generation model, which product teams across Yuanbao, WorkRally, Ima, and other services are helping refine through co-design. Tencent Cloud prices the model at USD 0.024 per 2K output image, and it supports text-to-image and image-to-image generation with up to five reference images. Tencent said an internal blind evaluation found it on par with ByteDance's Seedream 5.0 Pro and slightly better than Nano-Banana Pro and Qwen-Image-3.0 Pro.

Sep 22

Sep 22Tue
  1. ModelScopeAI score62

    inclusionAI open-sources Ming-Image-0.1-Design models for visual design

    AIinclusionAI open-sources the Ming-Image-0.1-Design family, two complementary 6B models for visual-design workflows, under an MIT License. Design generates complete UIs, dashboards, infographics, and posters up to 2048×2048 with native transparent RGBA output, and Layer decomposes flattened graphics into independently editable RGBA layers.

    Image from @ModelScope2022's post
  2. Fireworks AI BlogAI score65

    Fireworks releases Ember-1, a Kimi K3 variant that cuts reasoning tokens by about 40%

    AIFireworks Research released Ember-1, a specialized model built on Kimi K3 that it says delivers the same quality with 40% fewer tokens. Across five industry benchmarks, Ember-1 matched K3 max quality at a fraction of the cost, and in two customer A/B tests it used about 35% fewer tokens per task. It is available as a Research Preview on Serverless, and Fireworks is also launching training support for customized models.

    Why it matters: The source gives benchmark and A/B results for cutting reasoning tokens while holding quality, which bears on cost planning for coding and agent workloads.

  3. Tri DaoAI score44

    Rigel: 2.3B hybrid Mamba-2 MoE nears Llama-3.2-3B with <1% FLOPs

    AIMayank's Rigel, a 2.3B-parameter MoE (360M active) hybrid Mamba-2 model, was pretrained across H100, A100, V100 GPUs and TPU v5p/v6e on one codebase. The model lands within a few points of Llama-3.2-3B while using under 1% of its pretraining FLOPs. Tri Dao praised the work's engineering effort and the model's strength for its small size.

  4. Simon WillisonAI score60

    OpenAI launches GPT-6 Sol and Luna with 50% lower API prices

    AIOpenAI introduced GPT-6 Sol and GPT-6 Luna, which build on advances behind GPT-6 Astra. The company says it cut API prices 50% for Sol and Luna compared with GPT-5.6 promotional pricing, passing on caching and inference efficiency gains. Simon Willison notes GPT-6 Luna costs half of GPT-5.6 Luna and calls Luna his favorite model for building product features because of its cost and speed.

  5. Greg BrockmanAI score81

    OpenAI launches GPT-6 Sol and Luna with 50% lower API prices than GPT-5.6

    AIOpenAI introduced GPT-6 Sol and GPT-6 Luna, which it says bring much of the strength of GPT-6 Astra into faster and more affordable models. The company also reports more efficient caching and inference, with API prices 50% lower than GPT-5.6 promotional pricing.

    Why it matters: The quoted announcement names specific pricing and access changes for Sol and Luna, which matter for teams weighing cost against the Astra tier.

  6. Noam BrownAI score78

    OpenAI releases GPT-6 Sol and Luna at 50% lower API prices

    AIOpenAI has released GPT-6 Sol and GPT-6 Luna, which it says build on GPT-6 Astra and offer faster, more affordable performance. API prices for Sol and Luna are 50% lower than GPT-5.6 promotional pricing, and Luna now costs $0.10 input and $0.50 output per 1M tokens. The author also notes an earlier 80% Luna price cut at the end of July, with output dropping from $6 to $0.50 within two months.

    Why it matters: The source gives concrete API price cuts across two model tiers, making the cost trend across recent releases easy to track for developers.

  7. Mike KriegerAI score67

    Anthropic launches Claude Opus 5.5, leading in coding and knowledge work

    AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. According to the quoted launch post, it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. The author says it leads in coding and knowledge work and praises its writing quality.

    Why it matters: The quoted launch post gives a concrete cost comparison, useful for weighing Opus 5.5 against earlier Opus and Fable 5.1 models for routine work.

  8. AnthropicAI score71

    Anthropic releases Claude Opus 5.5, the first model in its Claude 5.5 family

    AIAnthropic has made Claude Opus 5.5 available today, introducing it as the first model in its new Claude 5.5 family. According to the quoted @claudeai post, it performs at the level of Claude Fable 5.1 on most tasks and costs 40% less to run than Opus 5.

    Why it matters: The post gives a concrete cost comparison against Opus 5 and names the model family, helping readers gauge the trade-off between price and performance.

  9. Black Forest Labs · new models on Hugging FaceAI score62

    Black Forest Labs releases FLUX 3 Action, a 7B open-weights robot world action model

    AIBlack Forest Labs released FLUX 3 Action, an open-weights 7B world action model that outputs robot joint commands from camera frames, robot state, and a text instruction. On the RoboLab-120 benchmark it reports 42.92% task success, ahead of Cosmos3-Nano-Policy at 36.8% and π0.5 at 28.0%. The model is fine-tuned on DROID, is distributed under the FLUX Kommunity License v.1.0, and runs in about 32 GB of GPU memory in bfloat16.

    Why it matters: The model card gives a benchmark comparison, parameter counts, and an action contract, so readers can judge how it compares with existing robot policies.

  10. Black Forest Labs · new models on Hugging FaceAI score58

    Black Forest Labs releases open-weights FLUX 3 Action SO-101 robot policy

    AIBlack Forest Labs has published FLUX 3 Action SO-101 on Hugging Face as an open-weights 7B world action model. It takes two camera frames, the robot state, and a text instruction, then returns the next 42 actions with predicted video frames, with 32 executed at 30 Hz before replanning. The card also provides a rank-32 LoRA fine-tuning recipe for user datasets and states that the application must enforce joint velocity, force, and workspace limits.

  11. Black Forest Labs · new models on Hugging FaceAI score60

    Black Forest Labs releases FLUX 3 Action base weights for robot adaptation

    AIBlack Forest Labs has released flux-3-action-base, an open-weights 7B world action model that takes camera frames, robot state, and a text instruction to output the next action chunk. The release is an adaptation component rather than a complete robot policy, and new embodiments require their own action heads. The source says the weights are paired with shared video VAE and Qwen3-VL-4B-Instruct text encoders and is governed by the FLUX Kommunity License v.1.0.

    Why it matters: The source separates the adaptation base from full robot policies and states the shared encoders and new-embodiment requirements, which clarifies what developers must still build for their robots.

  12. AI SupremacyAI score45

    TypeSafe AI's Jev Is a Non-LLM Probabilistic Classifier for Fast Software Decisions

    AITypeSafe AI released Jev, a transformer-based System-1 model that outputs calibrated probabilistic decisions instead of generating tokens, returning answers in 70–500 ms at $0.042 per million input tokens. The model is built for typed Choice, Score, and yes/no questions inside software pipelines, and it is available to everyone without a waitlist, with $5 in starting credits. Vercel, Cloudflare, LangChain, and Langfuse have added Jev to their platforms.

Sep 21

Sep 21Mon
  1. StepFunAI score58

    StepFun's Step 5 Preview scores 44 on Intelligence Index at lower cost

    AIStepFun's Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index at about $0.72 per task, matching Kimi K3 (max) at roughly 2.8x lower cost. The source reports strong reasoning results, including 46% on Humanity's Last Exam, but places it behind Qwen3.8 Max and GLM-5.3 (max) on agentic evaluations. Open weights are planned for October 15.

  2. Claude Apps Release NotesAI score62

    Anthropic launches Claude Opus 5.5, first model in its 5.5 family

    AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. The company says it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

    Why it matters: The release note gives a direct comparison to Claude Fable 5.1 and a 40% running cost reduction versus Opus 5, which helps readers weigh the tradeoff.

  3. Xiaomi MiMoAI score31

    Xiaomi's MiMo-V2.6-Pro reaches top 10 on Code Arena WebDev

    AIArena says Xiaomi's MiMo-V2.6-Pro debuted at about #10 overall on Code Arena: WebDev with a 1628-point AutoEval score, tying Claude Fable 5 (High). That is a 153-point gain over MiMo-V2.5-Pro's 1475, and it ranks about #3 among open-weights models under an MIT license. Arena notes the score is early, based on a reward model rather than live human votes, so rankings may shift as more votes arrive.

  4. Xiaomi MiMoAI score67

    Xiaomi MiMo open-sources Pro, Flash, and a 9B distilled model

    AIXiaomi MiMo announced open-source releases of Pro and Flash, the MiMo-V2.6-Distill-Qwen-9B model, a technical report, over 7K RL task environments, an end-to-end RL framework, and composable mini-harnesses. The attached table shows MiMo-V2.6-Distill-Qwen-9B after SFT and after RL compared with Qwen3.5-9B, with RL scores higher on most listed benchmarks, such as SWE-bench Verified at 66.2 versus 60.0.

    Why it matters: The table compares a 9B distilled model against Qwen3.5-9B on coding, cyber, and agent benchmarks, showing how the reinforcement learning stage changes results.

    Image from @XiaomiMiMo's post
  5. Xiaomi MiMoAI score78

    Xiaomi releases open-weight MiMo-V2.6 Pro and Flash omnimodal models

    AIXiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.

    Why it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

    Image from @XiaomiMiMo's post