The blind spot in one naturalness score
A study posted August 10 decomposes TTS “naturalness” into ten linguistically grounded perceptual dimensions. The authors built a dimension-level meta-evaluation benchmark from 860 utterances annotated by trained linguist raters and tested four MOS predictors and four Audio-LLM judges.
In the authors' results, MOS predictors largely collapsed onto acoustic signal quality. Audio-LLM judges detected selected problems but were prompt-sensitive and did not generalize across all dimensions. Neither group reliably captured the breadth of linguistically structured speech errors.
Product implication
Clean audio is not the same as correct words, prosody, rhythm or language switching. Multilingual voice agents should not rely on a single aggregate MOS. Teams can separately test pronunciation, semantic preservation, prosody, numbers, names and code-switching, with listening review by speakers of the target language.
Research boundary
This is a preprint, not completed peer review, covering 860 utterances and eight selected evaluators. Its results do not automatically generalize to every language, voice and synthesizer. It nevertheless demonstrates why an evaluator itself needs a meta-evaluation. The authors report releasing the data, schema and code.