NTH

Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry

AuthorsSzu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee

AffiliationsNational Taiwan University, Taipei, Taiwan · NVIDIA, Taiwan · Artificial Intelligence Center of Research Excellence (NTU AI-CoRE), NTU, Taiwan

October 2, 2026 2 min read
Watch on YouTube
The one-line take

The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.

Key results

+0.07
EER-alignment correlation

Pooled Spearman correlation between EER and human perceptual alignment across 35 model conditions.

-0.95
Effective-dimensionality correlation

Pooled Spearman correlation between effective dimensionality and perceptual alignment.

0.738
AM-Softmax bottleneck alignment

Perceptual alignment after reducing the ECAPA-TDNN embedding dimension to 3.

1.47%
ECAPA AAM-Softmax EER

EER for the ECAPA-TDNN AAM-Softmax model.

5994
VoxCeleb2-dev speakers

Number of speakers used to train the embedding models.

What the paper found

A NVIDIA-linked study challenges the standard use of equal error rate, or EER, as a proxy for human voice similarity in speech generation. Across speaker-embedding models trained on VoxCeleb2-dev with 5994 speakers, verification quality on VoxCeleb1-O showed almost no relationship to human judgments from VoxSim: the pooled Spearman correlation between EER and perceptual alignment was only +0.07. The decisive factor was the training objective and the geometry of the embedding space. Prototypical and Angular Prototypical losses consistently aligned better with listeners than classification losses such as AM-Softmax and AAM-Softmax, even without sacrificing EER; on ECAPA-TDNN, Angular Prototypical achieved a perceptual-alignment score of 0.406 versus 0.111 for AAM-Softmax, while both reached an EER of 1.47%. The proposed geometric measure, effective dimensionality, estimates how many directions speaker representations occupy. Across 35 model conditions, it correlated with perceptual alignment at −0.95, indicating that lower-dimensional speaker geometry better matches human perception. Explicit bottlenecking confirmed causality: reducing ECAPA-TDNN’s embedding dimension made the weakest 192-dimensional AM-Softmax model rise from 0.080 to 0.738 in perceptual alignment, although verification accuracy deteriorated. The results suggest that voice-similarity evaluation should prioritize embedding geometry over raw verification performance, with implications for TTS and voice-conversion systems built around models such as WavLM and ECAPA-TDNN.

Original abstract

Speaker verification (SV) models are commonly assumed to better capture nuances among speaker characteristics as verification accuracy improves, leading to their widespread use as automated proxies for human voice similarity in speech generation tasks. However, by establishing a human perceptual alignment metric and conducting systematic analysis, we demonstrate that perceptual alignment is governed far more by how a model is trained (its learning objective) than by how well it performs (EER). Notably, standard margin-based classification losses (e.g., AAM-Softmax) yield substantially lower perceptual alignment than prototypical metric losses, while EER itself fails to track human judgment, directly challenging the community's implicit assumption. We trace this divergence to embedding geometry, where a model's effective dimensionality ($d_{\mathrm{eff}}$) tracks perceptual alignment with a $-0.95$ rank correlation, revealing that the dimensional spread favored by classification losses fundamentally clashes with the low-dimensional nature of human voice perception. Imposing a dimensionality bottleneck compresses $d_{\mathrm{eff}}$ and raises perceptual alignment ($ρ_{\mathrm{align}}$) from 0.08 to 0.74, establishing a principled geometric criterion for evaluating voice similarity.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis
03Speech

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong, Sittipong Sripaisarnmongkol, Pakorn Nathong, Phatrasek Jirabovonvisut

A large voice-cloning model becomes a synthetic data generator for training a compact, reference-free Thai TTS system that runs on-device.

Read analysis