Skip to content
Read the original: Xiaomi MiMo· Published Pick68/100AI score68/100

Xiaomi releases MiMo-V2-TTS, a speech model with controllable emotion and singing

Original titleXiaomi MiMo-V2-TTS

AISummary

Xiaomi has launched MiMo-V2-TTS, a speech synthesis model that lets users describe the desired voice style in plain language. The model also supports dialects, character voices, non-verbal sounds such as coughs and sighs, and singing within one model. It was pretrained on over 100 million hours of speech data and refined with multi-dimensional reinforcement learning.

AIWhy it matters

The source gives concrete controls for emotion, dialect, singing, and non-verbal sounds, showing how a voice model can be directed through plain-language style prompts.

Read the original mimo.xiaomi.com

Source: Xiaomi MiMo · mimo.xiaomi.comPublished · added here