HeyGen Voice scores 83.1% overall on the Artificial Analysis Pronunciation Robustness benchmark, ranking #10 of 29 models.
Pronunciation Robustness measures whether Text to Speech models correctly pronounce challenging text across four categories, with human reviewers judging each clip against pre-agreed accepted pronunciations.
➤ Preserving Exact Sequences: 83.3%, ranking #2 behind SpaceXAI TTS at 85.7%
➤ Standalone Terms: 90.5%, with Eleven v3 Conversational leading at 95.1%
➤ Contextually Appropriate: 89.8%, with Gemini 3.8 Flash TTS leading at 97.9%
➤ Expanding Shorthand: 77.0%, compared to Qwen-Audio-3.1-TTS-Plus at 80.2%, with Eleven v4 leading at 94.1%
