Skip to contentSkip to stories

Updated

#Voice

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 20

Sep 20Sun
  1. OpenBMBAI score44

    MiniCPM-o Booking Desk: open-source real-time voice appointment agent built on MiniCPM-o 4.5

    AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.

Sep 18

Sep 18Fri
  1. Google AIAI score47

    Google's weekly recap: Gemini 3.8 Live, Dreambeans, CC, and more

    AIGoogle's weekly recap covers Gemini 3.8 Live and 3.8 Live Extended Thinking, described as its most advanced live dialogue audio models yet. It also notes Dreambeans, a GoogleLabs experiment curating daily personalized stories, is now generally available, and that CC has expanded into a shared agent for household coordination. Google Pics, a Workspace tool for generating and co-creating images, is now GA, alongside AlphaGenome Atlas, DeepMind's interactive genomics discovery platform.

Sep 17

Sep 17Thu
  1. xAI News (Grok)AI score42

    Grok Voice Transcribe 2.0 Doubles Accuracy of Predecessor at Same Price

    AIxAI released Grok Voice Transcribe 2.0, a speech-to-text model that is twice as accurate as Grok Voice Transcribe 1.0 at the same price, and ranks first for accuracy among 32 streaming models on the Artificial Analysis leaderboard. Batch transcription costs $0.10 per hour of audio and streaming $0.20 per hour, with diarization, timestamps, and key terms included. Existing Speech-to-Text API integrations gain the improvement with no code changes, and developers must pin grok-voice-transcribe-1.0 to stay on the older model during the transition.

Sep 16

Sep 16Wed
  1. inclusionAI (Ant Ling) · new models on Hugging FaceAI score55

    inclusionAI releases Realtime-Venus full-duplex audio-visual models on Hugging Face

    AIinclusionAI has published Realtime-Venus on Hugging Face with two 9B checkpoints: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for audio-only conversation. Both are built on MiniCPM-o 4.5 with a Qwen3-8B backbone and support full-duplex dialogue, proactive responses, and training-free long-video memory. The asynchronous Realtime-Venus-Harness runtime is hosted in a separate GitHub repository.

Sep 15

Sep 15Tue
  1. Google AI StudioAI score72

    Google releases Gemini 3.8 Live and 3.5 Transcribe for real-time voice apps

    AIGoogle AI Studio released Gemini 3.8 Live, a native speech-to-speech model with an Extended Thinking variant, and made it available through the Live API. Gemini 3.5 Transcribe, released last month, supports 85+ languages with a reported 4.0% streaming and 2.6% non-streaming Word Error Rate, and accepts a custom vocabulary of up to 1,000 terms. Live API audio pricing is listed at $0.005/min for input and $0.018/min for output.

    Why it matters: The post lists concrete Live API capabilities, per-minute audio pricing, and transcription accuracy figures, helping developers weigh voice agent options against their own cascaded pipelines.

  2. Google AI StudioAI score46

    Google launches Gemini 3.8 Live and Extended Thinking dialogue models

    AIGoogle introduced two live dialogue models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, available through AI Studio and the Gemini API. Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. The Extended Thinking variant targets high-complexity tasks with increased intelligence and multi-step reasoning.

  3. Google AIAI score72

    Google rolls out Gemini 3.8 Live and Extended Thinking across consumer, developer, and enterprise channels

    AIGoogle is rolling out Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking across several channels. Consumers get them in Search Live and Gemini Live, developers get public preview access through the Gemini API, and enterprises get private preview through Gemini Enterprise, with Customer Experience support coming soon.

    Why it matters: The post lays out where each Gemini 3.8 Live variant reaches consumers, developers, and enterprises, which clarifies access paths for a voice model release.

  4. Google AIAI score62

    Google releases Gemini 3.8 Live and 3.8 Live Extended Thinking audio models

    AIGoogle AI announces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking as its most advanced Gemini Audio models. Gemini 3.8 Live is built for scale, speed, and cost efficiency, handling mid-sentence interruptions, transitions across 97 languages, and visual context through Search Live. Gemini 3.8 Live Extended Thinking reasons and speaks in parallel, narrating its progress on multi-step tasks such as event planning.

  5. Google DeepMindAI score72

    Google DeepMind releases Gemini 3.8 Live models for real-time voice agents

    AIGoogle DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents. Extended Thinking scores 82.6 on Artificial Analysis' Speech to Speech Quality Index, 68.6% on τ-Voice, and 97.7% on Big Bench Audio. Gemini 3.8 Live is rolling out now in the Gemini API, Google AI Studio, and Search Live, with enterprise access in private preview.

    Why it matters: The release covers a voice model's benchmark results and availability across developer, enterprise, and consumer products, useful for judging voice agent options.

  6. Google AI StudioAI score72

    Google launches Gemini 3.8 Live and Extended Thinking voice models

    AIGoogle introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents that reason and speak simultaneously. The Extended Thinking version scores 82.6 on Artificial Analysis' Speech to Speech Quality Index and 97.7% on Big Bench Audio, while 3.8 Live targets scale and cost efficiency. Developers can access both through the Gemini API in Google AI Studio, and enterprise and consumer rollouts vary by product.

    Why it matters: The source names the two models, their access paths, and specific benchmark results, showing how the voice agent capabilities differ between the two tiers.

  7. Google DeepMindAI score33

    Gemini 3.8 Live Extended Thinking adds upgraded reasoning for real-time programming tutoring.

    AIGoogle DeepMind demonstrated 3.8 Live Extended Thinking acting as a programming tutor in Gemini Live. Both 3.8 Live models feature upgraded reasoning, near real-time visual understanding, automatic detection across 97 languages, and background tool calling that doesn't interrupt the chat. The Extended Thinking variant adds higher performance and precision for harder tasks and narrates its progress, and it is available in Gemini Live in the Gemini app or through the Gemini API via Google AI Studio.

  8. Google · Innovation & AIAI score52

    Google says its language technology now covers over 300 languages with new speech, data, and on-device tools

    AIGoogle reports that its technologies and products now power everyday interactions in more than 300 languages used by over 7 billion people, about 86% of the global population. The post describes new speech models, including Gemini 3.5 Live Translate and Gemini 3.5 Transcribe, plus the TranslateGemma open translation models trained across 55 languages.

  9. Gemini API ChangelogAI score62

    Google makes Gemini 3.8 Live models generally available for real-time voice

    AIGoogle has made two audio-to-audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, generally available through the Live API. Gemini 3.8 Live, model ID gemini-3.8-live, is the default for low-latency voice agents, with interleaved reasoning and asynchronous function calling. Gemini 3.8 Live Extended Thinking, model ID gemini-3.8-live-extended-thinking, supports background reasoning during live audio and is recommended when more reasoning is needed.

    Why it matters: The changelog names two model IDs and their intended use, showing how Live API developers can choose between low-latency voice and higher background reasoning.

Sep 11

Sep 11Fri

Sep 10

Sep 10Thu
  1. Greg BrockmanAI score72

    GPT-Live-1 becomes available in the OpenAI API for voice agents

    AIGPT-Live-1 is now available in the API, letting developers bring ChatGPT-style back-and-forth conversation into their apps. The quoted announcement says the voice agents can listen while they speak and can work with the models and harness developers choose.

    Why it matters: The quoted announcement describes a real-time voice model entering the API, which matters for builders weighing voice agents against existing stacks.

  2. Microsoft Foundry BlogAI score41

    Azure AI Speech LLM 2607 adds multilingual accuracy gains and phrase list customization

    AIMicrosoft released Azure AI Speech LLM 2607, which improves multilingual recognition, mixed-language audio handling, and domain-specific entity accuracy, and runs up to 3x faster than the previous 2605 release. A new dedicated phrase list parameter lets developers supply domain vocabulary, supporting 2,000+ entities, without embedding it in a prompt. The model is available through the Fast API and Real-Time API, is testable in the Foundry Playground, and is deployed automatically with no customer action required.

  3. Tencent HunyuanAI score60

    Tencent Hunyuan releases open-source AuK audio model for speech generation and editing

    AITencent Hunyuan has released AuK, an open-source foundation model for unified speech generation and editing that takes natural-language instructions and reference audio. It supports tasks including zero-shot TTS, timbre, style and emotion editing, denoising, and music separation. A companion AuK-Flash variant runs 4-step inference and is about 4.5 times faster under matched conditions, with code, weights, and a demo now available.

Sep 9

Sep 9Wed
  1. Microsoft Foundry BlogAI score62

    Microsoft Foundry's July and August 2026 updates bring Hosted Agents and Toolboxes to GA

    AIMicrosoft Foundry's July and August 2026 updates make Hosted Agents, Voice Live integration, and Toolboxes generally available. The post adds Claude tools on Azure, Model Router region and model pool changes, Foundry Local preview features, and updated Python, JavaScript, Java, and .NET SDK versions with migration notes.

    Why it matters: The roundup links each GA and preview change to code examples, migration notes, and runtime requirements, which helps developers judge what to upgrade and test first.

Sep 3

Sep 3Thu