Skip to contentSkip to stories

Updated

#Voice

Showing low-relevance items too. Hide low-relevance items

Sep 23

Sep 23Wed
  1. Google DeepMindAI score60

    Google DeepMind launches Gemini 3.8 Flash TTS and Flash-Lite TTS models

    AIGoogle DeepMind introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, text-to-speech models offering custom voice design, line-by-line performance control, and multilingual support across more than 100 languages. Flash TTS is rolling out to developers in the Gemini API and Google AI Studio and to everyone in Gemini Notebook, while Flash-Lite TTS is available to developers and in Google Vids. Voice replication requires consent verification, and generated audio carries SynthID watermarking.

    Why it matters: The source details the voice design, performance direction, and consent safeguards, showing how the model covers creative and high-volume use cases with access across several Google products.

  2. Google DeepMind · YouTubeAI score46

    Gemini 3.8 text-to-speech lets developers design and clone custom voices

    AIGoogle DeepMind's latest Gemini Audio models let developers design new vocal personas from natural language prompts, directing pacing, back channeling, and dialect shifts line by line. Developers can also recreate consistent adult voice profiles from a 30-second audio sample, with built-in consent verification, SynthID watermarking, and C2PA credentials.

  3. Baseten BlogAI score62

    Baseten launches NVIDIA Nemotron 3 Diarization with four latency profiles

    AIBaseten has made NVIDIA Nemotron 3 Diarization available as batch, streaming, and real-time diarized transcription presets. The single checkpoint serves four algorithmic latencies from 0.32 to 30.4 seconds, and the post reports DER of 9.8% on AISHELL-4 at the low profile versus 27.2% for Streaming Sortformer v2.1.

    Why it matters: The post shows one checkpoint serving four latency profiles with DER figures against named baselines, useful for judging real-time speaker labeling tradeoffs.

  4. ModelScopeAI score44

    NVIDIA releases Nemotron 3 Diarization for live speaker attribution

    AINVIDIA's Nemotron 3 Diarization is now available on ModelScope, labeling speakers and timestamps in streaming audio for up to eight speaker slots per conversation. The 99.2M-parameter model uses an end-to-end streaming architecture built on NVIDIA's Streaming Sortformer, running on Ampere, Hopper, and Blackwell GPUs via NeMo Speech C++. It is designed to pair with existing ASR systems such as Nemotron ASR, Parakeet, Canary, or Whisper to produce speaker-attributed transcripts.

    Image from @ModelScope2022's post

Sep 22

Sep 22Tue
  1. OpenBMBAI score59

    VoxWeft runs real-time interpretation locally on Apple Silicon using VoxCPM2

    AIOpenBMB highlights VoxWeft, an open-source simultaneous interpretation system for Apple Silicon built by developer @HenryZ30734018 on an MLX implementation of VoxCPM2. The system turns live speech into translated speech on-device, with first audio streaming in about 170 ms on an M5 MacBook. VoxCPM2 generates speech in 30 languages, supports direct language-pair interpretation, and clones a target voice from about 5 seconds of reference audio.

    Video from @OpenBMB's post
  2. Gemini API ChangelogAI score62

    Gemini 3.8 Flash TTS and Flash-Lite TTS become generally available with a new Voices endpoint

    AIGoogle made the Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS models generally available, along with the Gemini API Voices endpoint. Flash TTS is positioned for studio-grade voice fidelity and long-form multi-turn stability, while Flash-Lite TTS targets high-throughput, real-time voice agents and replaces gemini-3.1-flash-tts-preview. The update adds voice design, voice replication with consent verification, and access to 150+ prebuilt and custom voices.

Sep 21

Sep 21Mon
  1. Xiaomi MiMoAI score36

    Xiaomi MiMo-V2.6 unifies code, design, and tool use across creative outputs

    AIXiaomi's MiMo-V2.6 combines code, design, and tool use to build frontend interfaces, presentations, Figma-linked visual assets, and videos. The post says MiMo-V2.5-TTS supports narration in video production, and that the model can compose music, including an orchestral piece for around ten instruments that can be converted to MIDI. On Design Arena, the Pro version reportedly performs comparably to Claude Opus 5 and GPT-5.6 Sol.

    Image from @XiaomiMiMo's post
  2. Gemini NotebookAI score34

    Live Chat rolls out to all Ultra users on mobile

    AIGoogle's Gemini Notebook says Live Chat is now fully rolled out to all Ultra subscribers on mobile. The feature enables real-time voice conversations with notebooks in about 100 languages, powered by the latest audio models. Users can ask questions about their sources and receive step-by-step guidance largely hands-free.

    Video from @Gemini_Notebook's post

Sep 20

Sep 20Sun
  1. OpenBMBAI score44

    MiniCPM-o Booking Desk: open-source real-time voice appointment agent built on MiniCPM-o 4.5

    AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.

    Image from @OpenBMB's post

Sep 18

Sep 18Fri
  1. Google AIAI score47

    Google's weekly recap: Gemini 3.8 Live, Dreambeans, CC, and more

    AIGoogle's weekly recap covers Gemini 3.8 Live and 3.8 Live Extended Thinking, described as its most advanced live dialogue audio models yet. It also notes Dreambeans, a GoogleLabs experiment curating daily personalized stories, is now generally available, and that CC has expanded into a shared agent for household coordination. Google Pics, a Workspace tool for generating and co-creating images, is now GA, alongside AlphaGenome Atlas, DeepMind's interactive genomics discovery platform.

Sep 17

Sep 17Thu
  1. xAI News (Grok)AI score42

    Grok Voice Transcribe 2.0 Doubles Accuracy of Predecessor at Same Price

    AIxAI released Grok Voice Transcribe 2.0, a speech-to-text model that is twice as accurate as Grok Voice Transcribe 1.0 at the same price, and ranks first for accuracy among 32 streaming models on the Artificial Analysis leaderboard. Batch transcription costs $0.10 per hour of audio and streaming $0.20 per hour, with diarization, timestamps, and key terms included. Existing Speech-to-Text API integrations gain the improvement with no code changes, and developers must pin grok-voice-transcribe-1.0 to stay on the older model during the transition.

Sep 16

Sep 16Wed
  1. inclusionAI (Ant Ling) · new models on Hugging FaceAI score55

    inclusionAI releases Realtime-Venus full-duplex audio-visual models on Hugging Face

    AIinclusionAI has published Realtime-Venus on Hugging Face with two 9B checkpoints: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for audio-only conversation. Both are built on MiniCPM-o 4.5 with a Qwen3-8B backbone and support full-duplex dialogue, proactive responses, and training-free long-video memory. The asynchronous Realtime-Venus-Harness runtime is hosted in a separate GitHub repository.

Sep 15

Sep 15Tue
  1. Google AI StudioAI score72

    Google releases Gemini 3.8 Live and 3.5 Transcribe for real-time voice apps

    AIGoogle AI Studio released Gemini 3.8 Live, a native speech-to-speech model with an Extended Thinking variant, and made it available through the Live API. Gemini 3.5 Transcribe, released last month, supports 85+ languages with a reported 4.0% streaming and 2.6% non-streaming Word Error Rate, and accepts a custom vocabulary of up to 1,000 terms. Live API audio pricing is listed at $0.005/min for input and $0.018/min for output.

    Why it matters: The post lists concrete Live API capabilities, per-minute audio pricing, and transcription accuracy figures, helping developers weigh voice agent options against their own cascaded pipelines.

  2. Google AI StudioAI score46

    Google launches Gemini 3.8 Live and Extended Thinking dialogue models

    AIGoogle introduced two live dialogue models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, available through AI Studio and the Gemini API. Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. The Extended Thinking variant targets high-complexity tasks with increased intelligence and multi-step reasoning.

    Video from @GoogleAIStudio's post