Skip to contentSkip to stories

Updated

#Voice

Showing low-relevance items too. Hide low-relevance items

Sep 15

Sep 15Tue
  1. Google AIAI score72

    Google rolls out Gemini 3.8 Live and Extended Thinking across consumer, developer, and enterprise channels

    AIGoogle is rolling out Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking across several channels. Consumers get them in Search Live and Gemini Live, developers get public preview access through the Gemini API, and enterprises get private preview through Gemini Enterprise, with Customer Experience support coming soon.

    Why it matters: The post lays out where each Gemini 3.8 Live variant reaches consumers, developers, and enterprises, which clarifies access paths for a voice model release.

  2. Google AIAI score62

    Google releases Gemini 3.8 Live and 3.8 Live Extended Thinking audio models

    AIGoogle AI announces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking as its most advanced Gemini Audio models. Gemini 3.8 Live is built for scale, speed, and cost efficiency, handling mid-sentence interruptions, transitions across 97 languages, and visual context through Search Live. Gemini 3.8 Live Extended Thinking reasons and speaks in parallel, narrating its progress on multi-step tasks such as event planning.

    Video from @GoogleAI's post
  3. Google DeepMindAI score72

    Google DeepMind releases Gemini 3.8 Live models for real-time voice agents

    AIGoogle DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents. Extended Thinking scores 82.6 on Artificial Analysis' Speech to Speech Quality Index, 68.6% on τ-Voice, and 97.7% on Big Bench Audio. Gemini 3.8 Live is rolling out now in the Gemini API, Google AI Studio, and Search Live, with enterprise access in private preview.

    Why it matters: The release covers a voice model's benchmark results and availability across developer, enterprise, and consumer products, useful for judging voice agent options.

  4. Google AI StudioAI score72

    Google launches Gemini 3.8 Live and Extended Thinking voice models

    AIGoogle introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents that reason and speak simultaneously. The Extended Thinking version scores 82.6 on Artificial Analysis' Speech to Speech Quality Index and 97.7% on Big Bench Audio, while 3.8 Live targets scale and cost efficiency. Developers can access both through the Gemini API in Google AI Studio, and enterprise and consumer rollouts vary by product.

    Why it matters: The source names the two models, their access paths, and specific benchmark results, showing how the voice agent capabilities differ between the two tiers.

  5. Google DeepMindAI score33

    Gemini 3.8 Live Extended Thinking adds upgraded reasoning for real-time programming tutoring.

    AIGoogle DeepMind demonstrated 3.8 Live Extended Thinking acting as a programming tutor in Gemini Live. Both 3.8 Live models feature upgraded reasoning, near real-time visual understanding, automatic detection across 97 languages, and background tool calling that doesn't interrupt the chat. The Extended Thinking variant adds higher performance and precision for harder tasks and narrates its progress, and it is available in Gemini Live in the Gemini app or through the Gemini API via Google AI Studio.

    Video from @GoogleDeepMind's post
  6. Google · Innovation & AIAI score52

    Google says its language technology now covers over 300 languages with new speech, data, and on-device tools

    AIGoogle reports that its technologies and products now power everyday interactions in more than 300 languages used by over 7 billion people, about 86% of the global population. The post describes new speech models, including Gemini 3.5 Live Translate and Gemini 3.5 Transcribe, plus the TranslateGemma open translation models trained across 55 languages.

  7. Gemini API ChangelogAI score62

    Google makes Gemini 3.8 Live models generally available for real-time voice

    AIGoogle has made two audio-to-audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, generally available through the Live API. Gemini 3.8 Live, model ID gemini-3.8-live, is the default for low-latency voice agents, with interleaved reasoning and asynchronous function calling. Gemini 3.8 Live Extended Thinking, model ID gemini-3.8-live-extended-thinking, supports background reasoning during live audio and is recommended when more reasoning is needed.

    Why it matters: The changelog names two model IDs and their intended use, showing how Live API developers can choose between low-latency voice and higher background reasoning.

Sep 11

Sep 11Fri

Sep 10

Sep 10Thu
  1. Greg BrockmanAI score72

    GPT-Live-1 becomes available in the OpenAI API for voice agents

    AIGPT-Live-1 is now available in the API, letting developers bring ChatGPT-style back-and-forth conversation into their apps. The quoted announcement says the voice agents can listen while they speak and can work with the models and harness developers choose.

    Why it matters: The quoted announcement describes a real-time voice model entering the API, which matters for builders weighing voice agents against existing stacks.

  2. Microsoft Foundry BlogAI score41

    Azure AI Speech LLM 2607 adds multilingual accuracy gains and phrase list customization

    AIMicrosoft released Azure AI Speech LLM 2607, which improves multilingual recognition, mixed-language audio handling, and domain-specific entity accuracy, and runs up to 3x faster than the previous 2605 release. A new dedicated phrase list parameter lets developers supply domain vocabulary, supporting 2,000+ entities, without embedding it in a prompt. The model is available through the Fast API and Real-Time API, is testable in the Foundry Playground, and is deployed automatically with no customer action required.

  3. Tencent HyAI score60

    Tencent Hunyuan releases open-source AuK audio model for speech generation and editing

    AITencent Hunyuan has released AuK, an open-source foundation model for unified speech generation and editing that takes natural-language instructions and reference audio. It supports tasks including zero-shot TTS, timbre, style and emotion editing, denoising, and music separation. A companion AuK-Flash variant runs 4-step inference and is about 4.5 times faster under matched conditions, with code, weights, and a demo now available.

    Video from @TencentHunyuan's post

Sep 9

Sep 9Wed
  1. Microsoft Foundry BlogAI score62

    Microsoft Foundry's July and August 2026 updates bring Hosted Agents and Toolboxes to GA

    AIMicrosoft Foundry's July and August 2026 updates make Hosted Agents, Voice Live integration, and Toolboxes generally available. The post adds Claude tools on Azure, Model Router region and model pool changes, Foundry Local preview features, and updated Python, JavaScript, Java, and .NET SDK versions with migration notes.

    Why it matters: The roundup links each GA and preview change to code examples, migration notes, and runtime requirements, which helps developers judge what to upgrade and test first.

Sep 3

Sep 3Thu

Aug 29

Aug 29Sat
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score34

    Fun-ASR-Nano-2512 Gets vLLM-Native Packaging for Speech Transcription

    AIFunAudioLLM has released Fun-ASR-Nano-2512-vllm, a vLLM-native packaging of the official Fun-ASR-Nano-2512 checkpoint, with weights bitwise equal to the source and no new LoRA weights. The validated path runs on vLLM 0.27.1 with float32 through an OpenAI-compatible transcription endpoint, tested on one NVIDIA H100 80 GB GPU. The source-licensed model is Apache License 2.0, and other vLLM versions, accelerators, and quantizations require separate validation.

Aug 26

Aug 26Wed

Aug 24

Aug 24Mon

Aug 21

Aug 21Fri

Aug 3

Aug 3Mon
  1. Manus BlogAI score38

    Manus Adds ElevenLabs Connector for Chat-Based Audio Generation, Transcription, and Voice Apps

    AIManus has launched an ElevenLabs connector that lets users generate speech, transcribe recordings, clone voices, and build audio apps through a single chat. Users connect their authorized ElevenLabs account via Integrations, and audio is processed within their own ElevenLabs environment according to its policies. Availability depends on users having an active ElevenLabs account, with capabilities tied to their ElevenLabs plan and credit balance.

Jul 31

Jul 31Fri
  1. SkyworkAI score35

    Skywork AI Hardware Family's first Skywork Note batch sells out in one week

    AISkywork's first batch of its Skywork Note AI hardware device sold out one week after launch, prompting an accelerated rollout of the wider family, including the recording clip, the Recall pendant, and the TriRing AI ring. The company says the device is meant to capture real-world conversations and moments outside the screen, so users spend less time typing and more time away from it.

Jul 23

Jul 23Thu
  1. One Useful Thing (Ethan Mollick)AI score67

    Ethan Mollick's guide to choosing AI tools for agentic work

    AIEthan Mollick's guide says ChatGPT and Claude are the main choices for real work, since their agent modes can act on a computer. He separates agent modes that run on the company's computers from those that access the user's own computer. He recommends keeping approval settings on for sending, spending, or deleting, because of prompt injection risk. He also notes that Gemini currently lags for agentic work, though its Notebook and video tools are useful.

Jul 21

Jul 21Tue
  1. Andrej KarpathyAI score30

    Karpathy suggests long voice rambles help LLMs understand your intent

    AIAndrej Karpathy describes using /voice to ramble for about 10 minutes, sometimes as a short interview, to give an LLM context that would be tedious to type. He says LLMs reconstruct these messy streams of thought remarkably well, often returning a cleaner version than the speaker started with, which improves shared understanding and reduces later corrections.

Jul 15

Jul 15Wed

Jun 26

Jun 26Fri
  1. Qwen · new models on Hugging FaceAI score44

    Qwen3-ForcedAligner-0.6B-hf Adds Timestamp Alignment for Speech Transcripts

    AIQwen released Qwen3-ForcedAligner-0.6B-hf, a Transformers-format forced aligner that predicts timestamps for arbitrary units within up to 5 minutes of speech in 11 languages. The model accepts transcripts from any ASR system, and the documentation shows it paired with Qwen3-ASR-0.6B and NVIDIA Parakeet CTC. Until it ships in an official Transformers release, users must install Transformers from source.

Jun 18

Jun 18Thu
  1. Cohere · new models on Hugging FaceAI score43

    Cohere Releases Open-Source 2B Arabic Speech Recognition Model Transcribe Arabic

    AICohere and Cohere Labs released Cohere Transcribe Arabic, an open-source 2B-parameter Arabic automatic speech recognition model under Apache 2.0. It is optimized for Arabic, Arabic dialects, English, and Arabic-English code-switched speech, using a Conformer encoder-decoder architecture supported natively in Transformers. The model's average WER of 25.87 and CER of 11.80 on the Open Universal Arabic ASR Leaderboard, as of 07.07.2026, is reported in the source.

  2. Andrew NgAI score15

    DeepLearning.AI launches course on adding voice to AI agents

    AIDeepLearning.AI has launched a course, taught by VocalBridge CEO Ashwyn, on adding voice to AI agents and applications. It covers building voice agents that are both reliable and fast, with three projects: a voice-interactive game, an agent that gains a voice in about 10 lines of code, and an agent that places outbound calls via a make_phone_call function.

    Video from @AndrewYNg's post