Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
AuthorsJian Chen, You Zhang, Mark Vinton
AffiliationsDolby Laboratories
Resources
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.
Key results
Masked non-autoregressive decoding accelerates the text-to-semantic stage relative to IndexTTS2.
Tacit-TTS speaker-similarity score on LibriSpeech test-clean.
Word error rate on LibriSpeech test-clean.
Average WER for English targets using references from eight other languages.
Number of samples used to distill Tacit-TTS from IndexTTS2.
Tacit-TTS generates speech over 10× faster than IndexTTS2 for utterances longer than 5 seconds.
What the paper found
Tacit-TTS is a transcript-free, zero-shot voice-cloning system distilled from IndexTTS2 that replaces autoregressive text-to-semantic decoding with a masked, non-autoregressive Transformer. It estimates output length without a reference transcript by combining acoustic syllable peaks, detected with RMS energy and YIN, with target-text syllable counts, then preserves both discrete semantic codes and continuous latent features through dual-head distillation. A ReFlow-distilled flow-matching renderer reduces semantic-to-mel sampling from approximately 25 Euler steps to 4–8, while the masked decoder accelerates the T2S stage by 26.5×. On LibriSpeech test-clean, Tacit-TTS reaches 0.875 speaker similarity and 5.54 WER, and across cross-lingual references from eight languages it records 0.10% WER for English targets with no generation failures. The system remains effective for infant babble and synthetic gibberish, where transcript-dependent systems such as F5-TTS, MaskGCT, and CosyVoice2 can fail because ASR transcripts are unreliable. Using 874K distillation samples and running on a single NVIDIA A100 GPU, Tacit-TTS generates speech over 10× faster than IndexTTS2 for utterances longer than 5 seconds, while retaining comparable perceptual quality. The paper evaluates English and Mandarin datasets and uses OpenAI’s Whisper for English intelligibility scoring.
Original abstract
TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning further supports cross-lingual and non-lexical references. We validate this capability using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail due to unreliable ASR transcripts.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong, Sittipong Sripaisarnmongkol, Pakorn Nathong, Phatrasek Jirabovonvisut
A large voice-cloning model becomes a synthetic data generator for training a compact, reference-free Thai TTS system that runs on-device.