FireRedTTS3: Unified Speech Generation and Editing with Semantically Enriched Speech Representations
AuthorsFeiyu Shen, Kun Xie, Yichen Wu, Ziqi Dai, Yichen Han, Junjie Li, Xuelong Geng, Fenglong Xie, Lei Xie, Xu Tang, Yao Hu
FireRedTTS3 uses semantically enriched speech representations to make voice cloning, controllable speech generation, and audio editing more stable and capable.
Key results
Hours of diverse audio used to train the semantically enriched tokenizer
Languages supported by FireRedTTS3-Base
Chinese dialects supported by FireRedTTS3-Base
Average WER/CER achieved by FireRedTTS3-Base
Average speaker similarity achieved by FireRedTTS3-Base
Instruction-following score for acoustic-parameter specification
What the paper found
FireRedTTS3 addresses error accumulation in continuous autoregressive text-to-speech by introducing RedAE, a continuous speech tokenizer trained with semantic distillation from a frozen Audio Encoder. RedAE preserves fine acoustic detail without vector quantization, while its semantic supervision improves text-speech alignment and stabilizes an otherwise lightweight LLM-DiT generator. The system uses Qwen3-style Transformers and a Qwen3-1.7B backbone, with FireRedTTS3-Base targeting zero-shot voice cloning across 24 languages and 21 Chinese dialects, and FireRedTTS3-Instruct unifying voice cloning, natural-language voice design, and localized semantic or acoustic editing. RedAE was trained for 550k steps on 500k hours of diverse audio. On Seed-TTS-Eval, FireRedTTS3-Base reached an average error rate of 3.04% and speaker similarity of 78.8%, outperforming compared systems on the combined score. On MiniMax-MLS-Test, it achieved an average error rate of 3.75% and speaker similarity of 84.8% across 24 languages. For instruction-controlled voice design, evaluated with Gemini-2.5-pro, FireRedTTS3-Instruct scored 85.8% on Chinese acoustic-parameter specification, 82.0% on descriptive-style directives, and 69.7% on role-play, exceeding Qwen3-TTS-VD. The results show that semantically enriched continuous representations can support high-fidelity synthesis, multilingual cloning, controllable voice attributes, and precise speech editing without additional semantic branches or multi-stage tokenizer training.
Original abstract
Recent continuous autoregressive TTS models operate directly on continuous speech representations, preserving rich acoustic details while leveraging the instruction-following capabilities of text LLMs. This paradigm opens new possibilities for voice cloning, instruction-controlled voice design, and speech editing, but remains susceptible to error accumulation during autoregressive generation. Existing solutions often require additional semantic modules, multi-stage tokenizer training pipelines, or complex autoregressive architectures. In this work, we propose FireRedTTS3, a simple yet effective speech generation and editing framework that mitigates error accumulation at the representation level. Specifically, we leverage a frozen Audio Encoder trained on diverse speech understanding tasks as a semantic teacher to regularize the audio feature space. This improves text-speech alignment and stabilizes autoregressive generation while keeping the overall system simple. FireRedTTS3 provides two variants: FireRedTTS3-Base for multilingual and multi-dialect zero-shot voice cloning, and FireRedTTS3-Instruct for unified voice cloning, instruction-controlled voice design, and speech editing. Experiments show that FireRedTTS3-Base achieves the best average speech intelligibility and speaker similarity among compared systems on Seed-TTS-Eval and MiniMax-MLS-Test, while FireRedTTS3-Instruct outperforms competing systems on InstructTTSEval and Ming-Freeform-Audio-Edit. These results demonstrate that semantically enriched continuous speech representations, combined with a simple architecture, enable stable, controllable, and high-fidelity speech generation and editing. Code and models are available at https://github.com/FireRedTeam/FireRedTTS3.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.