Skip to contentSkip to stories

Updated

#Voice

Showing low-relevance items too. Hide low-relevance items

Jun 10

Jun 10Wed
  1. Xiaomi MiMoAI score67

    Xiaomi releases open-source MiMo Code V0.1 terminal coding assistant

    AIXiaomi MiMo has released MiMo Code V0.1, an open-source AI coding assistant for the terminal under the MIT license. It ships with MiMo V2.5, a multimodal model offered free for a limited time with a million-token context window. The tool automatically loads existing Claude Code skills, MCP servers and commands, and reuses API configuration, and it supports providers including Anthropic, OpenAI, DeepSeek, Kimi and GLM.

    Why it matters: The post specifies MiMo Code's Claude Code compatibility and MIT license, which bear directly on whether existing coding-agent setups can migrate without rework.

    Image from @XiaomiMiMo's post

Jun 5

Jun 5Fri

May 11

May 11Mon
  1. Soumith ChintalaAI score22

    Thinky previews real-time interaction models for human-AI collaboration

    AISoumith Chintala, a Thinky-linked voice, said the company is at step one of a plan to increase human-AI bandwidth and raise the ceiling of joint intelligence. He shared a preview of interaction models, described as real-time collaborative tools that talk, listen, watch, and think alongside people. A linked Thinking Machines post describes the approach and early results.

May 10

May 10Sun
  1. Thinking Machines LabAI score67

    Thinking Machines Lab previews interaction models for real-time human-AI collaboration

    AIThinking Machines Lab announced a research preview of interaction models that take in audio, video, and text continuously and respond in real time without external turn-detection harnesses. The model, TML-Interaction-Small, is a 276B-parameter MoE with 12B active parameters, paired with an asynchronous background model for sustained reasoning and tool use. The post reports competitive intelligence scores and lower turn-taking latency against GPT-realtime and Gemini Live models, along with new interactivity benchmarks where baseline models largely failed.

    Why it matters: The post explains a time-aligned, full-duplex design and benchmarks against turn-based models, showing how interaction and background reasoning can be split across two cooperating models.

May 1

May 1Fri

Mar 26

Mar 26Thu
  1. Guillaume Lample @ NeurIPS 2024AI score62

    Mistral releases Voxtral TTS, its first open-weight speech model

    AIMistral's Voxtral TTS is its first speech model, presented as an open-weight text-to-speech model that reportedly delivers SOTA performance at significantly lower cost with very low latency. It combines autoregressive generation of semantic speech tokens with flow-matching for acoustic tokens, and a technical report on its training methodology is being released.

    Image from @GuillaumeLample's post

Mar 17

Mar 17Tue
  1. Xiaomi MiMoAI score71

    Xiaomi releases MiMo-V2-Omni, an omni-modal model for agentic tasks

    AIXiaomi introduces MiMo-V2-Omni, a single model that fuses image, video, and audio encoders into a shared backbone with native tool calling and UI grounding. The company reports benchmark results against Gemini 3 Pro, Claude Opus 4.6, and GPT 5.2, and demonstrates browser-based shopping and video-publishing workflows run through the OpenClaw agent scaffold. It also states the model supports over 10 hours of continuous audio understanding.

    Why it matters: The page gives benchmark comparisons, a driving-risk demo, and browser-task walkthroughs, letting readers check how far the omni-modal claims extend into agent use.

Mar 13

Mar 13Fri
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score44

    Fun-CineForge Releases Open-Source Dubbing Pipeline, Model, and CineDub-CN Dataset

    AIFun-CineForge, from FunAudioLLM, is an open-source toolkit with an end-to-end dataset pipeline and an MLLM-based model for zero-shot movie dubbing across diverse cinematic scenes. The team built CineDub-CN, described as the first large-scale Chinese television dubbing dataset, and reports that its model outperforms state-of-the-art methods on audio quality, lip-sync, timbre transition, and instruction following. Inference code and checkpoints were released on March 16, 2026, and the model runs on a consumer-grade GPU.

Feb 19

Feb 19Thu

Feb 4

Feb 4Wed
  1. Guillaume Lample @ NeurIPS 2024AI score62

    Mistral's Voxtral Realtime streams speech with sub-200ms latency and open weights

    AIVoxtral Realtime is a natively streaming speech model for voice agents and live applications, with latency configurable down to sub-200ms. At 480ms it stays within 1-2% WER of the offline model, and the weights are released under Apache 2.0. The attached FLEURS chart compares word error rates across latency settings for ten languages, including Chinese.

    Image from @GuillaumeLample's post
  2. Guillaume Lample @ NeurIPS 2024AI score62

    Mistral releases Voxtral 2 transcription models with real-time option

    AIMistral announces Voxtral 2 with two transcription models: Voxtral Realtime, released under an Apache 2 license with latency configurable to sub-200 ms, and Voxtral Mini Transcribe 2, which adds speaker diarization, word-level timestamps, and context biasing. The models support 13 languages and are available through the Mistral API, which the post describes as one of the most cost-effective transcription APIs on the market. The attached chart shows word error rates on FLEURS across Italian, Spanish, English, German, Portuguese, French, Russian, Dutch, and Chinese at several latency settings.

    Image from @GuillaumeLample's post

Feb 2

Feb 2Mon

Jan 21

Jan 21Wed
  1. Mistral AI · new models on Hugging FaceAI score65

    Mistral releases open-weight Voxtral Mini 4B Realtime 2602 speech model

    AIMistral AI released Voxtral Mini 4B Realtime 2602, a multilingual realtime speech-transcription model with 13 supported languages under the Apache 2.0 license. The model has a configurable transcription delay from 240ms to 2.4s, and it matches leading offline open-source models at a 480ms delay. The source says it is optimized for on-device deployment and is currently supported only in vLLM.

    Why it matters: The source specifies the 480ms delay operating point, 4B size, Apache 2.0 license, and vLLM serving path, which matter for teams weighing realtime transcription deployment.

Jan 14

Jan 14Wed
  1. Chip HuyenAI score14

    Agentic Hackathon projects tackle long-running tasks, retrieval, and multimodal agents

    AIChip Huyen praised projects at last weekend's Agentic Hackathon, which hosted by MongoDB and Cerebral Valley, where she served as a judge. Teams tackled long-running tasks such as memory management, recovery from mid-task failures, and consistency across steps and sub-agents, along with adaptive retrieval across databases, search indices, and websites. Finalist demos are scheduled in San Francisco tomorrow, with talks by Douglas Eck.

    Image from @chipro's post

Dec 22, 2025

Dec 22, 2025Mon
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score58

    Alibaba's FunAudioLLM releases Fun-Audio-Chat-8B for low-latency voice interaction

    AIFunAudioLLM has released Fun-Audio-Chat-8B, a roughly 8B-parameter large audio language model for natural, low-latency voice interaction, under Apache 2.0. It uses Dual-Resolution Speech Representations with a 5Hz frame rate, which the source says reduces GPU hours by nearly 50%, and it supports English and Chinese.

Dec 14, 2025

Dec 14, 2025Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score38

    Alibaba Releases Fun-ASR-MLT-Nano-2512, an 800M Multilingual Speech Recognition Model

    AIAlibaba's FunAudioLLM released Fun-ASR-MLT-Nano-2512, an 800M-parameter multilingual speech recognition checkpoint on Hugging Face that supports 31 languages, with emphasis on East and Southeast Asian languages. It is trained on hundreds of thousands of hours of speech and is available through the FunASR toolkit. The source's benchmark tables cover the Fun-ASR family rather than this checkpoint, so no checkpoint-specific accuracy figures are reported.

  2. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score36

    Fun-ASR-Nano-2512 Speech Recognition Model Released by Tongyi Lab on Hugging Face

    AITongyi Lab has released Fun-ASR-Nano-2512, an end-to-end speech recognition large model trained on tens of millions of hours of real speech, supporting low-latency real-time transcription across 31 languages. The model, which has 800M parameters, targets industry use such as education and finance and claims 93% accuracy in far-field, high-noise conditions. It is available on Hugging Face and works with the FunASR toolkit.

Dec 10, 2025

Dec 10, 2025Wed
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score42

    Fun-CosyVoice3-0.5B-2512 Released as Open-Source Multilingual Text-to-Speech Model

    AIAlibaba's FunAudioLLM has released Fun-CosyVoice3-0.5B-2512, a 0.5B-parameter LLM-based text-to-speech model on Hugging Face, with an RL variant also published. The model supports zero-shot voice cloning across 9 languages and 18+ Chinese dialects and accents, with streaming output at latency as low as 150ms. On the source's test-en benchmark, it reports a 2.24% WER and 71.8% speaker similarity, and the RL version reports 1.68% WER.