Xiaomi releases open-source MiMo-V2.5-ASR speech recognition model with dialect support
Overview
Xiaomi MiMo has released MiMo-V2.5-ASR, an open-source speech recognition model that the company says achieves state-of-the-art results across multiple benchmarks.
Xiaomi states the model supports bilingual Chinese–English recognition, Chinese dialects including Wu, Cantonese, Hokkien, and Sichuanese, code-switching, and lyrics transcription.
According to Xiaomi, the model is also designed to handle noisy environments and multi-speaker conversations. The benchmark performance and robustness claims come from Xiaomi's own announcement and have not been independently verified in the reports.
Written by AI from the articles below · updated Oct 8, 9:12 PM ET
Check the sources:
Article timeline
Follow the coverage from different perspectives. Times are ET.
- Xiaomi MiMoPickXiaomi releases MiMo-V2.5-TTS series of speech synthesis models
AIXiaomi released the MiMo-V2.5-TTS Series, three speech synthesis models for stock voices, voice design, and voice cloning. The models accept natural-language style instructions and inline audio tags, and the source says the three models are free of charge for a limited time on the Xiaomi MiMo API platform. Xiaomi also open-sourced integration Skills for agent applications on GitHub.
- Xiaomi MiMoXiaomi releases open-source MiMo-V2.5-ASR speech recognition model with dialect support
AIXiaomi MiMo has released MiMo-V2.5-ASR, an open-source speech recognition model that the company says achieves state-of-the-art results across multiple benchmarks. The model supports bilingual Chinese–English recognition, Chinese dialects such as Wu, Cantonese, Hokkien, and Sichuanese, code-switching, and lyrics transcription. It is also designed to handle noisy environments and multi-speaker conversations.
Heat trend
Not enough continuous observations to show a trend yet.