Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions
AuthorsOluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols
Resources
This paper tests whether automated TTS judges can recognize the specific linguistic flaws that humans hear, revealing that today’s evaluators often reduce speech quality to acoustics or depend heavily on prompting.
Key results
Final balanced dataset size.
Linguistically grounded evaluation attributes.
UTMOSv2, DNSMOS-Pro, NISQA, and Audiobox-Aesthetics.
Kendall’s τ for Gemini 3.5 Flash with transcript and schema-guided per-dimension scoring.
Fraction of intentionally injected failures missed by majority human annotation.
What the paper found
“Beyond Naturalness” argues that TTS quality cannot be reduced to one holistic naturalness score. The researchers build the first dimension-level benchmark with 860 utterances annotated by trained linguists across 10 linguistically grounded dimensions covering word accuracy, prosody, and paralinguistic behavior. Speech is generated with Cartesia Sonic-3 from Harvard Sentences, EmergentTTS-Eval, and synthetic text, while GPT-5 and Praat-based manipulation inject controlled phoneme, stress, intonation, timing, emotion, speaker-consistency, and artifact errors. The audit compares 4 MOS predictors, including UTMOSv2, DNSMOS-Pro, NISQA, and Audiobox-Aesthetics, with 4 Audio-LLM judges: Google’s Gemini 3 Flash and Gemini 3.5 Flash, Alibaba’s Qwen3 Omni, and Step-Audio-2-Mini. MOS predictors primarily detect signal-level degradation, reaching their strongest performance on human plausibility and speech rate while largely missing phonetic accuracy, lexical stress, and prosodic boundaries. Audio-LLMs show selective, prompt-dependent sensitivity rather than general linguistic competence; Gemini 3.5 Flash with a transcript and per-dimension schema reached Kendall’s τ of 0.514 for phonetic accuracy. Schema guidance can recover word-level sensitivity, but jointly scoring all dimensions often causes output collapse. Human raters missed 32.3% of intentionally injected failures, underscoring the difficulty of perceptual error construction. The study concludes that interpretable, dimension-aware evaluation is necessary for reliable TTS diagnostics.
Original abstract
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.