SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
AuthorsYu Zhang, Ruiqi Li, Changhao Pan, Ke Lei, Xiang Yin, Cheng Yang
Resources
SwanTale aims to generate expressive multi-speaker dialogue and rich audio scenes from either natural-language instructions or a short voice reference.
Key results
Approximate scale of the SwanData-Caption annotation mixture.
SwanTale's zero-shot score on SwanBench-Speech.
Overall human-rated score on SwanBench-Scene.
SwanBench-Caption score after removing Unified MoE.
SwanBench-Caption score using Qwen3.0-Instruct-32B.
What the paper found
SwanTale, a ByteDance system, unifies instruct and zero-shot generation for multi-speaker speech, environmental audio, music, and local sound effects. Its SwanData-Caption pipeline produces approximately 70M caption records with structured environment, speaker, and fine-grained content descriptions, supported by targeted synthetic data and quality auditing. The model combines SwanVAE, which represents 48 kHz audio as 96-dimensional latents at 25 Hz, with a flow-matching DiT, reward-conditioned quality control, Engram memory for recurring caption patterns, and Unified MoE routing inspired by DeepSeek-style expert specialization. Curriculum learning preserves reference-audio voice cloning while adding natural-language voice design, and GRPO post-training optimizes pronunciation, stability, and speaker control. The caption encoder uses Qwen3.0-Instruct-8B, while Gemini 2.5 Pro and Gemini 3.5 Flash judge instruction-following and complex audio generation. On SwanBench-Speech, SwanTale reaches 0.95 timbre consistency for monologue zero-shot synthesis and leads expressive metrics for both monologues and dialogue. On SwanBench-Scene, it achieves an overall Mean MOS of 4.22, exceeding systems including Qwen3-TTS and Seedance 2.0. In SwanBench-Caption ablations, removing Unified MoE reduces instruction accuracy from 3.39 to 3.02, while replacing the caption encoder with Qwen3.0-Instruct-32B raises it to 3.70. Remaining challenges include long scenes beyond two minutes, complex music transitions, and precise continuous control of emotion, pauses, and local effects.
Original abstract
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine-grained content, while the zero-shot task uses reference audio together with the same fine-grained content. We address these tasks from both the data and model sides. First, we propose SwanData-Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi-level captions. Then, we propose SwanTale, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks. We introduce SwanVAE to support high-quality multi-audio-modality generation. Then, we adopt reward-conditioned quality control and Engram conditioning, along with Unified MoE for multi-task and multi-audio-modality modeling. In addition, we use curriculum learning and GRPO post-training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi-speaker speech and audio. Demos can be found at https://swanaigc.github.io/\#swantale.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.