Natural-Sounding Is Not Enough: What Automated TTS Judges Miss

A benchmark of 860 utterances across ten perceptual dimensions finds that MOS predictors and Audio-LLM judges both miss broad classes of linguistic speech errors.

The blind spot in one naturalness score

A study posted August 10 decomposes TTS “naturalness” into ten linguistically grounded perceptual dimensions. The authors built a dimension-level meta-evaluation benchmark from 860 utterances annotated by trained linguist raters and tested four MOS predictors and four Audio-LLM judges.

In the authors' results, MOS predictors largely collapsed onto acoustic signal quality. Audio-LLM judges detected selected problems but were prompt-sensitive and did not generalize across all dimensions. Neither group reliably captured the breadth of linguistically structured speech errors.

Product implication

Clean audio is not the same as correct words, prosody, rhythm or language switching. Multilingual voice agents should not rely on a single aggregate MOS. Teams can separately test pronunciation, semantic preservation, prosody, numbers, names and code-switching, with listening review by speakers of the target language.

Research boundary

This is a preprint, not completed peer review, covering 860 utterances and eight selected evaluators. Its results do not automatically generalize to every language, voice and synthesizer. It nevertheless demonstrates why an evaluator itself needs a meta-evaluation. The authors report releasing the data, schema and code.

Primary source