Xiaomi releases MiMo-V2-TTS, a speech model with controllable emotion and singing
Original titleXiaomi MiMo-V2-TTS
AISummary
Xiaomi has launched MiMo-V2-TTS, a speech synthesis model that lets users describe the desired voice style in plain language. The model also supports dialects, character voices, non-verbal sounds such as coughs and sighs, and singing within one model. It was pretrained on over 100 million hours of speech data and refined with multi-dimensional reinforcement learning.
AIWhy it matters
The source gives concrete controls for emotion, dialect, singing, and non-verbal sounds, showing how a voice model can be directed through plain-language style prompts.
Source: Xiaomi MiMo · mimo.xiaomi.comPublished · added here