Benchmarking 15+ TTS systems on Indic speech, by ear
A human evaluation framework for text-to-speech across Indic languages and Indian-accented English — what raters score, how the sampling makes systems comparable, and a per-language ranking the team can act on.
- Systems benchmarked
- 15+
- Coverage
- Indic · Indian-accented English
- Scored on
- naturalness · pronunciation · prosody
- Output
- per-language ranking
समस्या
Choosing a TTS vendor for Indic languages on published numbers means choosing on English evidence. The published scores are computed on English test sets, and the automatic metrics that generalise across languages do not hear the things that make synthesised Indic speech sound wrong.
A single mean opinion score has the opposite problem: it averages across languages, so a system that is strong in Hindi and unusable in Bengali lands in the same place as one that is mediocre in both.
निर्णय
The framework starts from what raters are actually being asked — naturalness, pronunciation and code-mixing accuracy, and prosody over long-form text — scored as separate judgements, because a system can be fluent and still mispronounce every proper noun.
The rater workflow and the sampling were designed so that scores are comparable across systems rather than merely collected: the same text, in the same conditions, across 15+ commercial and open systems, spanning Indic languages and Indian-accented English.
परिणाम
The results are reported as a per-language ranking, so the choice of system can be made language by language rather than from one averaged figure.
That output is what the team uses to select a vendor, and the protocol is repeatable as new systems appear.