NTH

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

AuthorsKunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong, Sittipong Sripaisarnmongkol, Pakorn Nathong, Phatrasek Jirabovonvisut

AffiliationsWayu Research · Paxa Labs · Typhoon · Sep 2026 Technical Report

September 20, 2026 2 min read
Watch on YouTube
The one-line take

A large voice-cloning model becomes a synthetic data generator for training a compact, reference-free Thai TTS system that runs on-device.

Key results

82M
Student model size

Parameter count of the fixed-voice Wayu-Paxa-TTS-Edge model.

68.2%
Challenge-Set Keyword Accuracy

Accuracy on the 1,531-item Thai benchmark targeting names, rare words, code-switching, informal spelling, and long sentences.

3.7%
Thai CER

Character error rate of the final bilingual student on 500 Thai evaluation utterances.

91.4%
Pause precision

Prosody pause precision of the final student, exceeding the OmniVoice teacher’s 89.9%.

87.9%
Best-of-K teacher accuracy

Oracle exact keyword accuracy at K = 118 teacher samples, versus 72.8% at K = 1.

65.5%
Code-switch accuracy gain

Accuracy with pretrained initialization, compared with 22.8% when key student modules are trained from scratch.

What the paper found

This paper presents a low-resource Thai text-to-speech pipeline that converts a 15-second voice reference into a compact fixed-voice model without any real speaker-specific corpus. OmniVoice generates synthetic speech, TLTK supplies Thai phonemization and tone handling, and quality filtering checks pronunciation, pauses, speaking rate, and duration before training a Kokoro and StyleTTS2-based student. The resulting 82M-parameter Wayu-Paxa-TTS-Edge runs without reference audio and reaches 68.2% Challenge-Set Keyword Accuracy on 1,531 difficult Thai sentences, with 3.7% Thai CER and 1.1% English CER. Its pause precision is 91.4%, higher than the OmniVoice teacher’s 89.9%, while its intra-word pause rate is 1.4%, indicating improved prosodic stability. The central engineering finding is that rejection sampling preserves difficult-text coverage: simply scaling filtered data can reduce keyword accuracy, whereas resampling rejected examples improves both correctness and pause placement. Pretrained Kokoro initialization is also crucial, raising code-switch accuracy from 22.8% to 65.5%. The teacher itself has substantial recoverable variation: oracle best-of-K sampling increases exact keyword accuracy from 72.8% at K = 1 to 87.9% at K = 118. The student remains behind Gemini 3.1 Flash TTS on keyword accuracy, but inference-only frontend changes using DeepSeek-V4-Flash and expanded TLTK handling recover 1.1 percentage points without acoustic-model retraining, showing that Thai TTS errors often originate in text normalization and phoneme representation rather than synthesis alone.

Original abstract

In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programmable data source to turn a short voice reference (e.g., 15 seconds) into a compact fixed-voice student trained entirely on synthetic speech. This setting makes pipeline design consequential: teacher errors become training targets, while filtering failed generations can reduce coverage of difficult texts. Thai further introduces challenges from ambiguous word boundaries, lexical tone, names and loanwords, numeric verbalization, and Thai-English code-switching. We study how text preparation, synthetic generation, quality filtering, rejection sampling, and frontend choices affect the resulting student, and where teacher limitations remain. We evaluate CER, Challenge-Set Keyword Accuracy, Prosody Pause Accuracy, speaker similarity, and speaking rate. The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-device Thai TTS without reference audio. It achieves 68.2% Challenge-Set Keyword Accuracy (85.5% of Gemini 3.1) and 91.4% pause precision, outperforming its OmniVoice teacher (89.9%) and reaching 94.8% of Gemini 3.1. It also achieves the lowest pause-placement error and intra-word pause rates among the three systems, and 3.7% and 1.1% CER on Thai and English, respectively. We open-source the model and evaluation framework for Thai TTS development.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis