Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts
AuthorsYuyue Wang, Xihua Wang, Xin Cheng, Yijing Chen, Ruihua Song
Resources
This paper proposes a new way to generate combined speech and sound directly from natural language using an LLM-based planner and a new benchmark for composite audio.
Key results
Total training data pool across composite, sound, and speech scenarios.
Held-out composite examples used to form PlanAudio-Bench from AudioSet.
Composite clips included in the 1.27M training pool.
PlanAudio’s composite-scenario Fr’echet Audio Distance, improving over pipeline and unified baselines.
PlanAudio’s composite-scenario Word Error Rate, reported for speech-related evaluation.
PlanAudio is initialized from Qwen2.5-1.5B as the autoregressive LLM backbone.
What the paper found
Renmin University of China’s PlanAudio reframes audio generation as Free-Form-Text-Prompt-to-Unified-Audio synthesis, where a single unconstrained prompt must produce speech, sound, or tightly interleaved composites without external text rewriting. The key novelty is a unified autoregressive LLM backbone initialized from Qwen2.5-1.5B that bypasses a separate text encoder and inserts a semantic latent Chain-of-Thought stage: the model first predicts a compact continuous planning sequence z, supervised by segment-level embeddings from Audio Flamingo 3, then generates hierarchical AudioCraft codec tokens conditioned on both the prompt and z. To evaluate this setting, the authors build PlanAudio-Bench from AudioSet with decoupled annotations via Whisper and Gemini-2.5 Pro, using 4,500 held-out composite examples and a 1.27M-example training pool spanning 371k composite, 451k sound, and 354k speech clips. Across composite benchmarks, PlanAudio improves substantially over pipeline and unified baselines such as VoiceLDM and AudioLDM2 combinations, reducing FAD from 22.9–25.2 down to 8.52 and lowering WER to 0.41 while achieving the best human ratings for quality, temporal correctness, semantic alignment, and authenticity. Ablations show that semantic latent CoT outperforms explicit natural-language CoT and acoustic-token CoT, and that a constant multi-scenario curriculum is crucial to avoid catastrophic forgetting.
Original abstract
Audio generation has made significant progress, yet synthesizing unified audio where speech and sounds are naturally composited remains a challenge. Current methods either rely on disjoint pipelines, which fail to capture fine-grained interactions, or require structured inputs and external text rewriting, which limits the flexibility of free-form text prompts. In this paper, we introduce a new task: Free-Form-Text-Prompt-to-Unified-Audio generation, which aims to directly synthesize unified audio containing speech, sound, and their composites from unconstrained natural language. To address this task, we propose PlanAudio, a unified, autoregressive LLM-based framework. First, it simplifies the model architecture by leveraging intrinsic LLM reasoning capability instead of traditional text encoders. Second, it introduces a semantic latent chain-of-thought mechanism, an implicit planning mechanism that bridges high-level semantic understanding and low-level acoustic synthesis. Furthermore, we create PlanAudio-Bench, a specialized benchmark for evaluating composite audio scenarios. We perform evaluations in the scenarios of speech, sound, and their composites. The results demonstrate that PlanAudio generally outperforms the existing pipeline and unified baselines, while staying competitive with models designed for a single scenario. Our analysis further reveals the superiority of semantic latent CoT over other CoT mechanisms and highlights the importance of continuous multi-scenario training curricula.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.