Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
AuthorsSzu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
AffiliationsNational Taiwan University, Taipei, Taiwan · NVIDIA, Taiwan · Artificial Intelligence Center of Research Excellence (NTU AI-CoRE), NTU, Taiwan
Resources
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Key results
Pooled Spearman correlation between EER and human perceptual alignment across 35 model conditions.
Pooled Spearman correlation between effective dimensionality and perceptual alignment.
Perceptual alignment after reducing the ECAPA-TDNN embedding dimension to 3.
EER for the ECAPA-TDNN AAM-Softmax model.
Number of speakers used to train the embedding models.
What the paper found
A NVIDIA-linked study challenges the standard use of equal error rate, or EER, as a proxy for human voice similarity in speech generation. Across speaker-embedding models trained on VoxCeleb2-dev with 5994 speakers, verification quality on VoxCeleb1-O showed almost no relationship to human judgments from VoxSim: the pooled Spearman correlation between EER and perceptual alignment was only +0.07. The decisive factor was the training objective and the geometry of the embedding space. Prototypical and Angular Prototypical losses consistently aligned better with listeners than classification losses such as AM-Softmax and AAM-Softmax, even without sacrificing EER; on ECAPA-TDNN, Angular Prototypical achieved a perceptual-alignment score of 0.406 versus 0.111 for AAM-Softmax, while both reached an EER of 1.47%. The proposed geometric measure, effective dimensionality, estimates how many directions speaker representations occupy. Across 35 model conditions, it correlated with perceptual alignment at −0.95, indicating that lower-dimensional speaker geometry better matches human perception. Explicit bottlenecking confirmed causality: reducing ECAPA-TDNN’s embedding dimension made the weakest 192-dimensional AM-Softmax model rise from 0.080 to 0.738 in perceptual alignment, although verification accuracy deteriorated. The results suggest that voice-similarity evaluation should prioritize embedding geometry over raw verification performance, with implications for TTS and voice-conversion systems built around models such as WavLM and ECAPA-TDNN.
Original abstract
Speaker verification (SV) models are commonly assumed to better capture nuances among speaker characteristics as verification accuracy improves, leading to their widespread use as automated proxies for human voice similarity in speech generation tasks. However, by establishing a human perceptual alignment metric and conducting systematic analysis, we demonstrate that perceptual alignment is governed far more by how a model is trained (its learning objective) than by how well it performs (EER). Notably, standard margin-based classification losses (e.g., AAM-Softmax) yield substantially lower perceptual alignment than prototypical metric losses, while EER itself fails to track human judgment, directly challenging the community's implicit assumption. We trace this divergence to embedding geometry, where a model's effective dimensionality ($d_{\mathrm{eff}}$) tracks perceptual alignment with a $-0.95$ rank correlation, revealing that the dimensional spread favored by classification losses fundamentally clashes with the low-dimensional nature of human voice perception. Imposing a dimensionality bottleneck compresses $d_{\mathrm{eff}}$ and raises perceptual alignment ($ρ_{\mathrm{align}}$) from 0.08 to 0.74, establishing a principled geometric criterion for evaluating voice similarity.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.
Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong, Sittipong Sripaisarnmongkol, Pakorn Nathong, Phatrasek Jirabovonvisut
A large voice-cloning model becomes a synthetic data generator for training a compact, reference-free Thai TTS system that runs on-device.