NTH

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

AuthorsOluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols

August 18, 2026 2 min read
Watch on YouTube
The one-line take

This paper tests whether automated TTS judges can recognize the specific linguistic flaws that humans hear, revealing that today’s evaluators often reduce speech quality to acoustics or depend heavily on prompting.

Key results

860
Benchmark utterances

Final balanced dataset size.

10
Perceptual dimensions

Linguistically grounded evaluation attributes.

4
MOS predictors evaluated

UTMOSv2, DNSMOS-Pro, NISQA, and Audiobox-Aesthetics.

0.514
Best phonetic-accuracy correlation

Kendall’s τ for Gemini 3.5 Flash with transcript and schema-guided per-dimension scoring.

32.3%
Human failure miss rate

Fraction of intentionally injected failures missed by majority human annotation.

What the paper found

“Beyond Naturalness” argues that TTS quality cannot be reduced to one holistic naturalness score. The researchers build the first dimension-level benchmark with 860 utterances annotated by trained linguists across 10 linguistically grounded dimensions covering word accuracy, prosody, and paralinguistic behavior. Speech is generated with Cartesia Sonic-3 from Harvard Sentences, EmergentTTS-Eval, and synthetic text, while GPT-5 and Praat-based manipulation inject controlled phoneme, stress, intonation, timing, emotion, speaker-consistency, and artifact errors. The audit compares 4 MOS predictors, including UTMOSv2, DNSMOS-Pro, NISQA, and Audiobox-Aesthetics, with 4 Audio-LLM judges: Google’s Gemini 3 Flash and Gemini 3.5 Flash, Alibaba’s Qwen3 Omni, and Step-Audio-2-Mini. MOS predictors primarily detect signal-level degradation, reaching their strongest performance on human plausibility and speech rate while largely missing phonetic accuracy, lexical stress, and prosodic boundaries. Audio-LLMs show selective, prompt-dependent sensitivity rather than general linguistic competence; Gemini 3.5 Flash with a transcript and per-dimension schema reached Kendall’s τ of 0.514 for phonetic accuracy. Schema guidance can recover word-level sensitivity, but jointly scoring all dimensions often causes output collapse. Human raters missed 32.3% of intentionally injected failures, underscoring the difficulty of perceptual error construction. The study concludes that interpretable, dimension-aware evaluation is necessary for reliable TTS diagnostics.

Original abstract

Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis